Approaching the event horizon
Radio static, life-support machinery, a low cosmic field, and a spoken line share the same trajectory.
Research project / 2026
An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together.
The idea
Generative models can synthesize high-quality video with sound and speech, yet interaction with generated content remains largely passive. EchoWM makes the generated scene enterable: a reference observation, a structured scene description, and a user control signal become the starting state for a world that keeps moving, sounding, and responding.
Audio is part of the state
Listen to the same rollouts that you watch. Environmental sound, music, and speech are generated with the scene and remain synchronized as the camera and subjects move.
Radio static, life-support machinery, a low cosmic field, and a spoken line share the same trajectory.
Spoken voice, distant ambience, and movement remain synchronized through the generated scene.
Visible action and generated speech evolve together within the same audio-visual rollout.
Environmental sound follows the traveler, landscape, and changing viewpoint as one generated world.
Selected demonstrations
Six selected audio-visual rollouts show how the same camera-intent interface transfers across first-person motion, third-person subjects, and diverse generated worlds.
Every clip below includes native generated sound.
Capability chapters / selected runs
The gallery follows the project proof structure: first-person world coverage, third-person camera--subject control, native audio-visual generation, and multi-turn viewpoint continuation.
How it works
Discrete commands and continuous poses become one metric-scale relative 6-DoF trajectory across first- and third-person scenes.
The model jointly generates 720p video, environmental sound, music, and speech instead of adding audio after visual synthesis.
A complementary data engine, progressive training, and autoregressive post-training support continuous navigation and multi-turn continuation.
Evaluation
EchoWM combines trajectory following with high-quality visual generation and native audio. On WBench it leads the navigation split by average score, while pairwise user studies show a clear preference for EchoWM on overall quality and semantic following.
Comparison on the 158-case WBench Navigation split. Results are sorted by average.
| Model | Avg. | Quality | Setting | Interaction | Consistency | Physical |
|---|---|---|---|---|---|---|
| EchoWM | 81.7 | 81.5 | 79.4 | 87.2 | 89.8 | 70.6 |
| EchoWM-Flash | 81.0 | 81.1 | 88.3 | 87.9 | 77.5 | 70.1 |
| HiDream-O1-World | 80.9 | 81.0 | 82.2 | 80.0 | 88.0 | 73.3 |
| Alaya-EVOKE | 80.8 | 82.8 | 83.8 | 78.6 | 86.9 | 72.1 |
| LingBot-World (fast v2) | 79.4 | 81.8 | 76.8 | 82.8 | 86.5 | 69.1 |
| Kling 3.0 | 79.0 | 81.4 | 91.0 | 69.4 | 83.7 | 69.3 |
| LingBot-World (base-camera) | 78.5 | 78.9 | 72.6 | 80.1 | 89.9 | 71.2 |
| Wan 2.7 | 78.1 | 81.5 | 91.4 | 64.4 | 81.6 | 71.8 |
| HY-World 1.5 (AR-distill) | 78.1 | 78.1 | 72.2 | 86.8 | 86.9 | 66.3 |
| HY-Video 1.5 | 77.9 | 77.6 | 85.6 | 71.4 | 87.4 | 67.4 |
| LingBot-World (fast) | 77.4 | 79.4 | 77.9 | 79.2 | 84.9 | 65.7 |
| HappyOyster | 76.8 | 77.3 | 74.2 | 84.9 | 84.3 | 63.5 |
| Lyra 2.0 (4-step AR) | 76.4 | 77.1 | 73.2 | 85.6 | 79.3 | 66.7 |
| Seedance 1.5 | 76.2 | 82.1 | 82.9 | 66.3 | 81.3 | 68.4 |
| Cosmos3-Super | 76.0 | 76.4 | 91.6 | 58.0 | 85.0 | 69.2 |
| SANA-WM (4-step AR) | 76.0 | 79.3 | 76.1 | 82.2 | 80.7 | 61.9 |
| DreamX-World (5B AR) | 75.0 | 77.5 | 80.8 | 78.6 | 74.9 | 63.3 |
| ABot-World | 74.7 | 76.8 | 71.4 | 84.0 | 79.5 | 61.7 |
| Cosmos 2.5 | 74.6 | 72.9 | 83.3 | 63.1 | 86.5 | 67.4 |
| Cosmos3-Nano | 74.2 | 77.4 | 87.3 | 59.1 | 83.7 | 63.6 |
| LTX-2.3 | 74.2 | 77.1 | 85.2 | 66.4 | 77.2 | 64.9 |
| Genie 3 | 73.9 | 75.2 | 72.5 | 73.4 | 82.6 | 65.7 |
Aggregate judgments over 200 cases. “Both good” and “both bad” preserve the two distinct same-quality outcomes.
| Competitor | Dimension | Competitor preferred | Both good | Both bad | EchoWM preferred |
|---|---|---|---|---|---|
| LingBot-World-v2 | Semantic following | 15.14% | 24.93% | 25.93% | 34.00% |
| Initial-world consistency | 7.71% | 62.36% | 15.29% | 14.64% | |
| Motion | 20.00% | 13.71% | 41.00% | 25.29% | |
| Spatial-temporal consistency | 14.71% | 27.36% | 25.29% | 32.64% | |
| Visual aesthetics | 19.93% | 28.00% | 19.57% | 32.50% | |
| Overall preference | 21.57% | 4.21% | 27.71% | 46.50% | |
| HappyOyster | Semantic following | 29.44% | 51.50% | 5.44% | 13.63% |
| Initial-world consistency | 15.44% | 42.81% | 2.25% | 39.50% | |
| Motion | 24.56% | 11.88% | 44.63% | 18.94% | |
| Spatial-temporal consistency | 10.44% | 18.19% | 13.69% | 57.69% | |
| Visual aesthetics | 23.81% | 1.13% | 22.19% | 52.88% | |
| Overall preference | 27.06% | 2.19% | 7.63% | 63.13% |
Abstract
Generative models can synthesize high-quality video with sound and speech, yet interaction with generated content remains largely passive; interactive world models support continuous control but are typically silent or tied to specific embodiments and action spaces.
We introduce EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music, and speech.
We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data.
To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation.
Extensive evaluations show that EchoWM achieves strong trajectory following and high visual quality on public world-model benchmarks, generalizes across real, game, and stylized environments, supports first- and third-person interaction across varied subjects, and maintains synchronized environmental sound and speech over long-horizon generation.
From content to participation
EchoWM turns generative media from something watched into a world that can be entered, steered, and heard evolving over time.
Read the work@misc{echowm2026,
title={EchoWM: Open and Enterable Omnimodal World Models},
year={2026},
note={Research project},
}