arXiv Machine Learning By Pardis Taghavi, Tingyu Guo, Jonas Lossner, Gaurav Pandey, Reza Langari

ActionSplice: In-Flight Action Editing for Interactive World Models

Read the original on arXiv Machine Learning →

ActionSplice is an inference framework for chunk‑autoregressive video world models that allows in‑flight action editing without re‑sampling completed evaluations. It formulates the problem as Counterfactual State Transport (CST), using a lightweight corrector to move the backbone representation toward the state induced by a revised action at the same solver step. Two variants, CST*R and CST*T, update either the entire active chunk or only its suffix, achieving significant reductions in rollback‑relative LPIPS and providing speedups over waiting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 8

Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling

Hi-FLoop introduces a hierarchical state‑feedback framework for multi‑agent traffic simulation that reconciles decision time scales over an 8‑second rollout. The model uses eight scene‑level Worlds to maintain joint hypotheses, with an 8‑second Goal, 2‑second Preview, and 1‑second Control hierarchy, and commits only executed prefixes every 0.5 seconds to preserve factual consistency. A joint preview interaction graph and a prefix‑frozen A‑to‑B cascade enable sparse interaction refinement and accurate state recovery, achieving an overall score of 0.689987 on the H‑D public‑validation split and strong oracle‑minADE performance. whyItMatters":"The paper presents a novel multi‑timescale approach that improves consistency and realism in long‑horizon traffic simulations, as evidenced by its competitive evaluation metrics."

arXiv Computer Vision
Sep 4

Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

The paper introduces Statebench, a benchmark for evaluating how well video generators track world states across segments, focusing on past-visible, occluded-process, and complex-transition states. It also proposes Stateagent, a method that maintains an explicit entity-state representation, updates it with new prompts, and uses the resulting state to guide video continuation. Experiments show Stateagent raises the overall state score from 45.2 to 69.3 and improves one‑minute story generation.

By Yingmao Miao, Pengfei Zhang, Chaoran Xu, Meng Yu, Jing Tang, Xiangxiang Chu, Chao Shen, Chenhao Lin
arXiv AI
Jul 20

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

arXiv:2607. 15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires.

By Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada