Hugging Face Trending Papers

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

arXiv AI
6d ago

HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.

By Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
arXiv Computer Vision
Sep 11

World in World: Explore the World with World Models

World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.

By Chenxi Song, Yanming Yang, Chi Zhang
arXiv Computer Vision
Sep 18

Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation

The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.

By Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
Hugging Face Trending Papers
Aug 18

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.