Abstract
Video generation is moving beyond isolated clips toward persistent long-form narratives, where characters, voices, events, and sound must remain coherent across many shots. We present the long-video variant of JoyAI-Echo-1.5, a unified audio-visual generation system built around composable cross-shot memory. It aggregates visual evidence from multiple prior shots and speaker cues from speech-filtered full-shot audio, preserving appearance and voice identity while supporting flexible text, image, and memory conditioning. A causal few-step generator, trained with progressive audio-visual teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts, improves stability over extended generation. On a benchmark of 100 stories and 3,000 shots, JoyAI-Echo-1.5 achieves the best result on six of seven automatic metrics, covering cross-shot consistency, imaging quality, text alignment, and speech fidelity. Human evaluation further favors JoyAI-Echo-1.5 over HappyOyster across all five long-video dimensions, with the largest margin in audio-visual synchronization.
Key Conclusions
Identity persists across distant shots.
Composable visual and speech-filtered audio memory delivers the strongest ViCLIP, Self-CIDS, and speaker-consistency scores among all evaluated systems.
Consistency does not trade away instruction following.
JoyAI-Echo-1.5 reaches the best CLIP score of 0.2868 while preserving changing actions, scenes, camera views, and shot-specific narrative details.
Audio and video remain one coherent timeline.
Joint generation achieves 0.9674 speech recall and a 28.1-point human-preference lead over HappyOyster in audio-visual synchronization.
Few-step generation remains stable over long rollouts.
An 8-step distilled generator combines causal audio-visual teacher forcing with short- and long-horizon Self-Gradient Forcing on self-generated histories.
Results
Evaluation on 100 stories and 3,000 shots measures cross-shot consistency, quality, text alignment, speech fidelity, and human preference.
Long-horizon automatic evaluation
JoyAI-Echo-1.5 ranks first on six of seven metrics and second on aesthetic quality. All metrics are higher-is-better.
| Method | Cross-Shot Consistency | Video Quality | Text | Speech | |||
|---|---|---|---|---|---|---|---|
| ViCLIP | Self-CIDS | Voice | Aesthetic | Imaging | CLIP | Recall | |
| JavisDiT++ | 0.6955 | 0.5621 | 0.7933 | 0.5019 | 0.6084 | 0.2522 | 0.0951 |
| LTX-2.3 | 0.7047 | 0.6135 | 0.6847 | 0.5436 | 0.4751 | 0.2559 | 0.9304 |
| LTX-2.5 | 0.7775 | 0.6640 | 0.7603 | 0.6019 | 0.5804 | 0.2683 | 0.9545 |
| ShotStream + MMAudio | 0.7988 | 0.7308 | 0.7945 | 0.5255 | 0.6044 | 0.1998 | 0.0297 |
| StoryMem + MMAudio | 0.7897 | 0.7491 | 0.7643 | 0.5265 | 0.6320 | 0.2368 | 0.0301 |
| HappyOyster (Directing) | 0.7535 | 0.6940 | 0.7705 | 0.5606 | 0.6701 | 0.2544 | 0.1085 |
| JoyAI-Echo-1.0 | 0.8026 | 0.7793 | 0.8129 | 0.5679 | 0.7058 | 0.2658 | 0.9489 |
| JoyAI-Echo-1.5 | 0.8264 | 0.7937 | 0.8524 | 0.5705 | 0.7467 | 0.2868 | 0.9674 |
Human evaluation against HappyOyster
Good-Similar-Bad pairwise comparison for long-horizon audio-visual generation.
| Aspect | JoyAI-Echo-1.5 | Similar | HappyOyster |
|---|---|---|---|
| Story-instruction following | 38.6% | 43.8% | 17.7% |
| Dialogue/background-audio following | 39.2% | 37.3% | 23.5% |
| Memory/identity consistency | 24.8% | 59.5% | 15.7% |
| Audio-visual synchronization | 48.4% | 31.4% | 20.3% |
| Audio quality | 44.4% | 22.2% | 33.3% |
Authors
Echo Team
BibTeX
@techreport{joyai2026echo15,
title = {Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds},
author = {{Echo Team @ Joy Future Academy, JD}},
institution = {Joy Future Academy, JD},
year = {2026},
month = {August}
}