arXiv Machine Learning

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

arXiv AI
Sep 21

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

The study examines how different vision‑language‑action (VLA) policies execute a manipulation task by comparing the geometry of their end‑effectors across 15,000 closed‑loop LIBERO rollouts. By pairing 3,600 configuration‑matched policy executions, the authors find that when both policies succeed, their end‑effector trajectories are much closer (median DTW distance 0.0120 m) than when only one succeeds (0.0380 m), a pattern consistent across all tasks, policy pairs, and nine representations. Even successful executions remain as far from same‑task demonstrations as the demonstrations are from each other, indicating that task‑associated geometry, rather than training data overlap, drives these differences.

By Xingyu Lin, Zhuang Li, Zhongrun Wu, Shouquan Zhou, Dehui Du
arXiv AI
6d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal