Video Generation
2 September 2026
H3 Video Model Found to Have World Model Capabilities
N
New-ZZZ desk
X @Hailuo_AI · 2 weeks ago
The Hailuo AI team created H3 as a video generation model, but experiments showed that its capabilities extend beyond the original task. The model was able to use its built-in understanding of language to control characters and the camera, exhibiting signs of a so-called world model—a system capable of accounting for the structure and dynamics of a scene.
The adaptation required only about 8,000 examples and training just 0.199% of H3’s parameters. The authors see this result as an example of how open models enable researchers to uncover hidden capabilities in existing systems.
Why it matters
- —The result shows that video models can acquire a more general understanding of scenes without full retraining.
- —Adapting a small fraction of parameters can unlock new ways to control characters and the camera.
- —Open access helps third-party teams discover capabilities not anticipated by the model's creators.
Key facts
- H3 was originally developed for video generation.
- Around 8,000 training examples were used for the experiment.
- The changes affected only 0.199% of the model's trainable parameters.
- H3's language understanding was directly adapted to control characters and the camera.
The full text is in the original source. Here we provide a brief summary and key facts.