Technical Report | August 2026 | Echo Team @ Joy Future Academy, JD

JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories

A long-video system that preserves character appearance, speaker identity, narrative continuity, and synchronized sound across multi-shot stories.

10 min+ long-horizon generation
A/V Memory cross-shot identity & voice

Abstract

Video generation is moving beyond isolated clips toward persistent long-form narratives, where characters, voices, events, and sound must remain coherent across many shots. We present the long-video variant of JoyAI-Echo-1.5, a unified audio-visual generation system built around composable cross-shot memory. It aggregates visual evidence from multiple prior shots and speaker cues from speech-filtered full-shot audio, preserving appearance and voice identity while supporting flexible text, image, and memory conditioning. A causal few-step generator, trained with progressive audio-visual teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts, improves stability over extended generation. On a benchmark of 100 stories and 3,000 shots, JoyAI-Echo-1.5 achieves the best result on six of seven automatic metrics, covering cross-shot consistency, imaging quality, text alignment, and speech fidelity. Human evaluation further favors JoyAI-Echo-1.5 over HappyOyster across all five long-video dimensions, with the largest margin in audio-visual synchronization.

Key Conclusions

01

Identity persists across distant shots.

Composable visual and speech-filtered audio memory delivers the strongest ViCLIP, Self-CIDS, and speaker-consistency scores among all evaluated systems.

02

Consistency does not trade away instruction following.

JoyAI-Echo-1.5 reaches the best CLIP score of 0.2868 while preserving changing actions, scenes, camera views, and shot-specific narrative details.

03

Audio and video remain one coherent timeline.

Joint generation achieves 0.9674 speech recall and a 28.1-point human-preference lead over HappyOyster in audio-visual synchronization.

04

Few-step generation remains stable over long rollouts.

An 8-step distilled generator combines causal audio-visual teacher forcing with short- and long-horizon Self-Gradient Forcing on self-generated histories.

Results

Evaluation on 100 stories and 3,000 shots measures cross-shot consistency, quality, text alignment, speech fidelity, and human preference.

Long-horizon automatic evaluation

JoyAI-Echo-1.5 ranks first on six of seven metrics and second on aesthetic quality. All metrics are higher-is-better.

Method Cross-Shot Consistency Video Quality Text Speech
ViCLIP Self-CIDS Voice Aesthetic Imaging CLIP Recall
JavisDiT++ 0.69550.56210.79330.50190.60840.25220.0951
LTX-2.3 0.70470.61350.68470.54360.47510.25590.9304
LTX-2.5 0.77750.66400.76030.60190.58040.26830.9545
ShotStream + MMAudio 0.79880.73080.79450.52550.60440.19980.0297
StoryMem + MMAudio 0.78970.74910.76430.52650.63200.23680.0301
HappyOyster (Directing) 0.75350.69400.77050.56060.67010.25440.1085
JoyAI-Echo-1.0 0.80260.77930.81290.56790.70580.26580.9489
JoyAI-Echo-1.5 0.82640.79370.85240.57050.74670.28680.9674

Human evaluation against HappyOyster

Good-Similar-Bad pairwise comparison for long-horizon audio-visual generation.

Aspect JoyAI-Echo-1.5 Similar HappyOyster
Story-instruction following38.6%43.8%17.7%
Dialogue/background-audio following39.2%37.3%23.5%
Memory/identity consistency24.8%59.5%15.7%
Audio-visual synchronization48.4%31.4%20.3%
Audio quality44.4%22.2%33.3%

Authors

Echo Team

BibTeX

@techreport{joyai2026echo15,
  title        = {Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds},
  author       = {{Echo Team @ Joy Future Academy, JD}},
  institution  = {Joy Future Academy, JD},
  year         = {2026},
  month        = {August}
}