Curate what matters
Shot detection, motion-aware segmentation, technical quality checks, and multimodal review retain complete physical events.
Technical report ↗
StrucPhysVideo generates high-fidelity video while learning the regularities beneath it—persistence, contact, motion, and state change.
The complete set of 21 selected generations across physical phenomena, everyday scenes, robot manipulation, and action-conditioned prediction.
StrucPhysVideo connects interaction-rich data, grounded descriptions, and action conditioning in a single Physical AI program.
Shot detection, motion-aware segmentation, technical quality checks, and multimodal review retain complete physical events.
Structured labels ground objects, actions, camera motion, and temporally ordered state transitions in visible evidence.
Text-image and end-effector action conditions guide generation while preserving visual quality and temporal continuity.
Physics-IQ Verified Score
+2.8 pts over the next-ranked model
* indicates our reproduction. Scores for the other comparison models are taken from the Physics-IQ benchmark snapshot dated 16 Sep 2026. Verified Score combines spatial overlap, spatiotemporal overlap, weighted spatial overlap, and normalized pixel error. Higher is better.
Relative scores normalized to the bidirectional baseline; higher is uniformly better.
Evaluation note The LingBot-Video result is reproduced by our team; the other comparison-model scores reflect the Physics-IQ Verified benchmark snapshot dated 16 Sep 2026. AgiBot results are controlled training-distribution diagnostics under each model’s native conditioning interface; they do not establish held-out generalization or isolate architecture alone.