arXiv Computer Vision By Di Wen, Ruodi Zhang, Kailun Yang, Kunyu Peng

Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models

Read the original on arXiv Computer Vision →

The paper examines latent action models that encode transitions between video frames using algebraic constraints such as additive composition and antisymmetric reversal. It demonstrates that these algebraic consistency conditions do not reliably certify temporal structure, as unconstrained models can achieve similar error reductions and constrained models still outperform unconstrained ones even after temporal pairings are destroyed. The authors find that preserving temporal pairing offers no consistent advantage on downstream tasks and that a direct repair objective yields only marginal improvement, recommending a more rigorous validation protocol.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 17

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.

By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv AI
Sep 10

ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.

By Zhengshu Zhang, Zhiyuan Li