Video world model for Physical AI

StrucPhysVideo
moves like the world.

StrucPhysVideo generates high-fidelity video while learning the regularities beneath it—persistence, contact, motion, and state change.

object persistence
continuous motion
45.5%Physics-IQ Verified
#1 / 17models evaluated
TI2V + IA2Vlanguage and action conditioned
01 — Generated worlds

See physical intelligence
unfold over time.

The complete set of 21 selected generations across physical phenomena, everyday scenes, robot manipulation, and action-conditioned prediction.

02 — How it works

From observable interaction
to predictive world model.

StrucPhysVideo connects interaction-rich data, grounded descriptions, and action conditioning in a single Physical AI program.

01

Curate what matters

Shot detection, motion-aware segmentation, technical quality checks, and multimodal review retain complete physical events.

02

Describe what changes

Structured labels ground objects, actions, camera motion, and temporally ordered state transitions in visible evidence.

03

Predict what follows

Text-image and end-effector action conditions guide generation while preserving visual quality and temporal continuity.

TI2VText + image → video
IA2VImage + action → video
03 — Quantitative results

Leading physical
prediction.

45.5%

Physics-IQ Verified Score
+2.8 pts over the next-ranked model

Physics-IQ Verified

Image-to-video leaderboard

Snapshot · 16 Sep 2026
StrucPhysVideo45.5
Cosmos3-Super Image2Video42.7
LingBot-Video*40.4
MiniMax H3 (FL2VA)39.8
Cosmos3-Nano37.3
MiniMax H3 Max36.2
Grok Imagine Video34.8
Magi-1 24B + GeoPhys (BoN; OP)33.7
Hunyuan Video 1.533.4
Cosmos3-Edge32.7
Wan 2.2 14B32.2
CogVideoX-5B31.8
Kandinsky-WM 1.030.8
Magi-1 24B (OP)30.2
Wan 2.2 5B27.7
Sora 226.5
P-Video25.3

* indicates our reproduction. Scores for the other comparison models are taken from the Physics-IQ benchmark snapshot dated 16 Sep 2026. Verified Score combines spatial overlap, spatiotemporal overlap, weighted spatial overlap, and normalized pixel error. Higher is better.

Physics caption ablation

Grounded supervision compounds.

33.2%Wan2.2-5B+8.4 pts vs. reproduced baseline
45.5%LingBot-Video-30B+7.6 pts vs. reproduced baseline
IA2V · AgiBot paired diagnostic

Higher frame fidelity, action conditioned.

PSNR ↑23.89
SSIM ↑0.8721
LPIPS ↓0.0880
Three-model comparison on eight fixed training-distribution clips, 256 × 256, frames 1–20.
DROID · 253-frame rollout

Long-horizon action following

Relative scores normalized to the bidirectional baseline; higher is uniformly better.

Trajectory3.88×
Depth3.52×
Interaction1.83×
LPIPS⁻¹3.30×

Evaluation note The LingBot-Video result is reproduced by our team; the other comparison-model scores reflect the Physics-IQ Verified benchmark snapshot dated 16 Sep 2026. AgiBot results are controlled training-distribution diagnostics under each model’s native conditioning interface; they do not establish held-out generalization or isolate architecture alone.

StrucPhysVideo

Generate the world.
Understand its rules.

Back to top