The paper audits frozen decoder‑only large language models (LLMs) on geometric reasoning tasks using parametric CAD constraints. It probes hidden states for linear decodability, forced‑choice generation, activation‑level influence, and behavioral steerability, finding that pretraining improves decoding of local geometric relations but not sketch‑level DOF status. The study shows that decodable information is not always actionable: generation often fails to express it, and steering interventions do not reliably control outputs, revealing divergences among decodability, generation, activation influence, and steerability.
By Man Liang, Xinzhao Cheng, Faizan Wajid
The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.
By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
The study examines how different vision‑language‑action (VLA) policies execute a manipulation task by comparing the geometry of their end‑effectors across 15,000 closed‑loop LIBERO rollouts. By pairing 3,600 configuration‑matched policy executions, the authors find that when both policies succeed, their end‑effector trajectories are much closer (median DTW distance 0.0120 m) than when only one succeeds (0.0380 m), a pattern consistent across all tasks, policy pairs, and nine representations. Even successful executions remain as far from same‑task demonstrations as the demonstrations are from each other, indicating that task‑associated geometry, rather than training data overlap, drives these differences.
By Xingyu Lin, Zhuang Li, Zhongrun Wu, Shouquan Zhou, Dehui Du
arXiv:2609.39971v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
The paper introduces Interaction‑Aligned Pruning (IAprune), a training‑free method for visual token pruning in embodied manipulation tasks. IAprune jointly decides per‑frame budget and token selection, using semantic‑motion spatial agreement to choose between conservative and aggressive coverage, and applies geometric residual correction to focus on under‑represented boundaries. Experiments on four policies, three simulation benchmarks, and a real‑robot platform show that IAprune matches unpruned performance on LIBERO while achieving up to 1.54× speed‑up and 1.48× acceleration on a real robot.
By Jintao Cheng, Weibin Li, Haozhe Wang, Gang Wang, Yipu Zhang, Xiaoyu Tang, Jin Wu, Xieyuanli Chen, Yunhui Liu, Wei Zhang
The paper investigates how large language models (LLMs) share a common Fisher‑Rao geometry in their next‑token probability distributions, revealing that behaviour largely determines this geometry while activation geometry depends on coordinate choices. Across transformer, state‑space, and recurrent architectures, output geometries align more closely than activation geometries, and this shared structure facilitates semantic‑category transfer and improves agreement with human word choices as models scale and train. The study further demonstrates that geometry can guide minimum‑disturbance interventions, enabling reusable control that preserves behaviour better than Euclidean methods and enhances steering, editing, attribution, dictionary learning, and fine‑tuning.
By Dario Picozzi