VLANeXt: Recipes for Building Strong VLA Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper presents VLANeXt, a Vision‑Language‑Action (VLA) model built from a unified framework that systematically analyzes design choices across foundational components, perception essentials, and action modeling. By dissecting these dimensions, the authors identify 12 key findings that form a practical recipe for strong VLA models, and demonstrate that VLANeXt outperforms state‑of‑the‑art methods on LIBERO and LIBERO‑plus benchmarks while also excelling in real‑world experiments. The study further extends VLANeXt into a family of variants—scaled models, latent‑action pretraining, latent predictive representation learning, and world action modeling—showing that the core recipe remains effective across different scales and emerging paradigms.
arXiv:2601. 03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining significant attention for their promising generalization capabilities.
StarVLA-α is a streamlined Vision‑Language‑Action (VLA) model that reduces architectural and pipeline complexity to facilitate systematic study of VLA design choices. By employing a strong VLM backbone and minimal design, it achieves competitive performance across multiple benchmarks (LIBERO, SimplerEnv, RoboTwin, RoboCasa) and outperforms the baseline π₀.₅ by 20% on the RoboChallenge benchmark. The authors plan to release the code to support future VLA research.
arXiv:2606. 19297v1 Announce Type: new Abstract: Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain after adaptation.
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain after adaptation. Failures on knowledge-sensitive tasks are ambiguous, conflating missing knowledge with poor generalization of low-level control.
arXiv:2609.36118v1 Announce Type: new Abstract: Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone la...