arXiv Computer Vision By Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag

HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Read the original on arXiv Computer Vision →

HiPhy introduces a hierarchical reinforcement learning framework for video generation that enforces physical laws at both local and global levels. It addresses the challenge of multi-principle interactions—such as buoyancy and fluid dynamics occurring simultaneously—by ensuring each principle’s temporal dynamics and the overall scene’s coherence. The authors also provide a 50K-prompt dataset and the MultiPhyBench benchmark, demonstrating that HiPhy outperforms existing methods, especially in scenes with multiple concurrent physical principles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 21

CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

CompAdapt is a physics-consistent text-to-video generation framework that extends diffusion-based models to handle composite physical behaviors such as coupled motions, multi-stage transitions, and multi-object collisions. It translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial parameters. The system introduces dynamics-aware prior matching for one-shot adaptation to new physical environments and a physics-aware latent feature fusion module to enhance visual fidelity during fast, complex motion, outperforming existing physics-constrained baselines on physics-focused T2V benchmarks.

By Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology, China), Yukang Ding (Taobao, Alibaba Group, China), Hui Li (Harbin Institute of Technology, China), Wangmeng Zuo (Harbin Institute of Technology, China)
Hugging Face Trending Papers
Jul 21

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions.

arXiv Computer Vision
Sep 4

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World is a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation without external offline modules. It jointly models physics, geometry, and appearance as native world states and uses a unified Omni-Camera representation to support diverse tasks and flexible motions. The framework also propagates physical dynamics across future frames, couples appearance and geometry in a single generative process, and scales to complex scenarios with the Puffin-16M dataset of 15 million vision‑language‑camera triplets and 1 million trajectories.

By Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy