Research project / 2026

EchoWMOpen and Enterable Omnimodal World Models

An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together.

Songchun Zhang1,3,*, Yaowei Li2,3,*, Junhao Zhuang3,*, Weiyang Jin3,4,*, Haoyu Wang5, Xin Lu6, Shiyi Zhang5, Haoran Li3, Xiaoxiao Ma3,6, Yumin Li3,4, Yijun Liu5, Yaofeng Su7, Yanwen Ma8, Haoyu Wu3,4, Zihan Su5, Yue Ma1, Lvmin Zhang9, Haoyang Huang3, Zeyue Xue3,‡, Anyi Rao1,†, Nan Duan3,†

1HKUST 2PKU 3JD Future Academy, JD 4HKU 5THU 6USTC 7FDU 8Beihang University 9Stanford University

*These authors contributed equally. Corresponding authors. Project lead.

Native audio-visual output First-person + third-person Multi-turn continuation
VIDEO / AUDIO / CONTROL Scroll to enter

The idea

Content that responds.

Generative models can synthesize high-quality video with sound and speech, yet interaction with generated content remains largely passive. EchoWM makes the generated scene enterable: a reference observation, a structured scene description, and a user control signal become the starting state for a world that keeps moving, sounding, and responding.

Audio is part of the state

Sound stays inside the world.

Listen to the same rollouts that you watch. Environmental sound, music, and speech are generated with the scene and remain synchronized as the camera and subjects move.

01 / SPEECH

Approaching the event horizon

Radio static, life-support machinery, a low cosmic field, and a spoken line share the same trajectory.

02 / SPEECH

Voice across the snowfield

Spoken voice, distant ambience, and movement remain synchronized through the generated scene.

03 / ON-SCREEN VOICE

Speech inside the scene

Visible action and generated speech evolve together within the same audio-visual rollout.

04 / ENVIRONMENT

Traveler at the turquoise shore

Environmental sound follows the traveler, landscape, and changing viewpoint as one generated world.

Native synchronized video + environmental sound + speech

Selected demonstrations

Enter from any viewpoint.

Six selected audio-visual rollouts show how the same camera-intent interface transfers across first-person motion, third-person subjects, and diverse generated worlds.

01 / TPP

Paper flight over the Forbidden City

Camera motion follows a lightweight subject across a large architectural environment.

02 / TPP

Robot at the jungle portal

Subject motion and camera evolution remain coherent inside a dense stylized world.

03 / TPP

Continuous camera--subject control

A long rollout demonstrates coordinated viewpoint and subject evolution.

04 / TPP

WBench navigation case 108

Continuous trajectory following is maintained over a long generated sequence.

05 / TPP

Bioluminescent forest

Camera and subject motion remain coordinated through a dense, luminous environment.

06 / FPP

Above the ridge

High-speed observer motion preserves a persistent visual and acoustic world.

07 / TPP + SPEECH

The room left behind

Creaking timber, broken-window wind, a toy engine, and a whispered question.

08 / TPP + SPEECH

Alien jungle signal

Shallow water, glowing flora, a crashed ship, and a voice inside the mist.

09 / TPP

Biomechanical rider

Metallic hoof impacts, distant explosions, fire, and low machinery hum.

10 / TPP

Bioluminescent flight

Wingbeats move through a resonant forest of glowing trees and drifting mist.

11 / TPP + SPEECH

Inside the bloodstream

Thrusters cut through biological flow, organic resonance, and a pilot's murmur.

12 / TPP + SPEECH

Train at the void

Steam, locomotive rumble, cosmic wind, and an engineer approaching the unknown.

13 / TPP

Cyberpunk threshold

Forward motion carries the scene and its urban energy through a luminous portal.

14 / TPP + SPEECH

Desert passage

Soft footfalls and sparse wind lead toward a structure buried in the cliffs.

15 / TPP + SPEECH

Above the palace

Wind, distant birds, soft chimes, and a child's voice cross the sunset skyline.

16 / TPP

Mushroom route

Engine revs and tires on dirt animate a saturated forest at oversized scale.

17 / TPP + SPEECH

Crystal sky islands

Wind threads through crystal formations while wings echo across open air.

18 / TPP

Snowball mountain

A remote snowy world with motion, atmosphere, and sound generated as one scene.

19 / FPP

Temple compass

First-person travel through leaves, bird calls, and dripping jungle moisture.

Capability chapters / selected runs

Four ways to enter a world.

The gallery follows the project proof structure: first-person world coverage, third-person camera--subject control, native audio-visual generation, and multi-turn viewpoint continuation.

How it works

One intent.
Many worlds.

01

Camera intent

Discrete commands and continuous poses become one metric-scale relative 6-DoF trajectory across first- and third-person scenes.

02

Omnimodal response

The model jointly generates 720p video, environmental sound, music, and speech instead of adding audio after visual synthesis.

03

Long-horizon world

A complementary data engine, progressive training, and autoregressive post-training support continuous navigation and multi-turn continuation.

Evaluation

Strong control.
Synchronized worlds.

EchoWM combines trajectory following with high-quality visual generation and native audio. On WBench it leads the navigation split by average score, while pairwise user studies show a clear preference for EchoWM on overall quality and semantic following.

81.7WBench average
87.2interaction score
46.50%overall preference vs. LingBot
63.13%overall preference vs. HappyOyster
TABLE 01

WBench navigation

Comparison on the 158-case WBench Navigation split. Results are sorted by average.

Comparison on the 158-case WBench Navigation split.
ModelAvg.QualitySettingInteractionConsistencyPhysical
EchoWM81.781.579.487.289.870.6
EchoWM-Flash81.081.188.387.977.570.1
HiDream-O1-World80.981.082.280.088.073.3
Alaya-EVOKE80.882.883.878.686.972.1
LingBot-World (fast v2)79.481.876.882.886.569.1
Kling 3.079.081.491.069.483.769.3
LingBot-World (base-camera)78.578.972.680.189.971.2
Wan 2.778.181.591.464.481.671.8
HY-World 1.5 (AR-distill)78.178.172.286.886.966.3
HY-Video 1.577.977.685.671.487.467.4
LingBot-World (fast)77.479.477.979.284.965.7
HappyOyster76.877.374.284.984.363.5
Lyra 2.0 (4-step AR)76.477.173.285.679.366.7
Seedance 1.576.282.182.966.381.368.4
Cosmos3-Super76.076.491.658.085.069.2
SANA-WM (4-step AR)76.079.376.182.280.761.9
DreamX-World (5B AR)75.077.580.878.674.963.3
ABot-World74.776.871.484.079.561.7
Cosmos 2.574.672.983.363.186.567.4
Cosmos3-Nano74.277.487.359.183.763.6
LTX-2.374.277.185.266.477.264.9
Genie 373.975.272.573.482.665.7
TABLE 02

Pairwise user study

Aggregate judgments over 200 cases. “Both good” and “both bad” preserve the two distinct same-quality outcomes.

Aggregate pairwise user-study judgments over 200 cases. Each cell reports the percentage of judgments.
CompetitorDimensionCompetitor preferredBoth goodBoth badEchoWM preferred
LingBot-World-v2Semantic following15.14%24.93%25.93%34.00%
Initial-world consistency7.71%62.36%15.29%14.64%
Motion20.00%13.71%41.00%25.29%
Spatial-temporal consistency14.71%27.36%25.29%32.64%
Visual aesthetics19.93%28.00%19.57%32.50%
Overall preference21.57%4.21%27.71%46.50%
HappyOysterSemantic following29.44%51.50%5.44%13.63%
Initial-world consistency15.44%42.81%2.25%39.50%
Motion24.56%11.88%44.63%18.94%
Spatial-temporal consistency10.44%18.19%13.69%57.69%
Visual aesthetics23.81%1.13%22.19%52.88%
Overall preference27.06%2.19%7.63%63.13%

Abstract

An enterable model for generative media.

Generative models can synthesize high-quality video with sound and speech, yet interaction with generated content remains largely passive; interactive world models support continuous control but are typically silent or tied to specific embodiments and action spaces.

We introduce EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music, and speech.

We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data.

To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation.

Extensive evaluations show that EchoWM achieves strong trajectory following and high visual quality on public world-model benchmarks, generalizes across real, game, and stylized environments, supports first- and third-person interaction across varied subjects, and maintains synchronized environmental sound and speech over long-horizon generation.

From content to participation

Watch less.
Enter more.

EchoWM turns generative media from something watched into a world that can be entered, steered, and heard evolving over time.

Read the work

EchoWM

Open and Enterable Omnimodal World Models

Research project / 2026

BibTeX
@misc{echowm2026,
  title={EchoWM: Open and Enterable Omnimodal World Models},
  year={2026},
  note={Research project},
}