arXiv:2609.36416v2 Announce Type: replace-cross
Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isol...
By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam
UniTrackPLA introduces a unified panorama-language-action model that simultaneously handles instruction‑guided navigation and dynamic person tracking for embodied robots. Its Panoramic‑Aware Encoding preserves azimuthal and temporal structure, allowing a shared vision‑language backbone to generate continuous waypoint chunks for both tasks. The model also employs World‑Action Consistency to predict future visual states and verify waypoint prefixes, enabling reliable action reuse and replanning when inconsistencies arise. A new OmniTrackNav‑Bench dataset and extensive real‑world experiments demonstrate significant performance gains over prior methods.
By Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang
LangMap introduces a human‑verified benchmark for language‑conditioned goal navigation (LGN) that spans four hierarchical semantic levels—scene, room, region, and instance—within real‑world indoor 3D scans. The dataset, built on HM3D, contains 18K tasks with concise and detailed descriptions for 414 object categories, and its contrastive annotation protocol ensures high‑quality, discriminative region and instance labels. Evaluation shows that LangMap’s descriptions improve text‑to‑view matching accuracy by 23 points over GOAT‑Bench and achieve a 92.5% unique‑and‑correct match rate in an independent human audit, while a proposed RGB‑only baseline, PlaNaVid, attains top‑tier success rates without depth or 3D scene representations.
By Bo Miao, Weijia Liu, Jun Luo, Lachlan Shinnick, Jian Liu, Thomas Hamilton-Smith, Yuhe Yang, Zijie Wu, Vanja Videnovic, Feras Dayoub, Anton van den Hengel
The paper introduces AEGIS, a buffer‑free, layer‑wise orthogonal gradient projection framework that enables continuous fine‑tuning of Vision‑Language Models for robotic manipulation while protecting pre‑trained visual reasoning. AEGIS uses pre‑training activation statistics as static anchors and applies a Wasserstein‑2 transport penalty to generate an anchor‑restoration gradient, followed by a dual‑backward pass that orthogonalizes task gradients against this restoration vector. Experiments on the LIBERO manipulation benchmark with PaliGemma2‑3B show that AEGIS preserves Visual Question Answering performance and baseline loss while achieving continuous action convergence, all without replay buffers, teacher models, or co‑training data.
By Guransh Singh
The paper presents a matched‑budget audit framework for evaluating recaptioned image‑text supervision distributions, which are created by applying a captioning policy, captioner, and source corpus to generate dense descriptions for text‑to‑image training. The framework operates under a fixed text budget of 64 tokens and produces a five‑axis profile covering prompt coverage, image‑conditioned faithfulness, and caption surface quality, using controllable basic units as a common claim metric. The authors apply the framework to seven paired comparisons across five public corpora, demonstrating improvements in supported claim units and revealing consistent long‑vs‑dense trade‑offs across different judges and budgets, and they release the audited multi‑source recap corpus and audit artifacts.
By Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu
arXiv:2610. 01871v1 Announce Type: cross Abstract: Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge.
By Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
Reconstruct, Practice, Go Real (RPG) is a framework that enables autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities from offline data, creates simulation practice tasks, and uses feedback to diagnose failures, develop new symbolic skills, refine existing ones, and revise the system prompt. Across 22 manipulation tasks, RPG raises task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming baselines and achieving perfect success on 30 physical trials after calibration.
By Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
REALM is a retrospective knowledge distillation framework that enables causal decoding of behavior from local field potentials (LFPs). It trains a bidirectional Mamba‑2 teacher on multi‑session data using continuous masked autoencoding, then distills its representations into a compact causal student model. The resulting LFP‑only decoder achieves the highest mean accuracy among compared methods, surpassing state‑of‑the‑art baselines while using fewer parameters and less pretraining time.
By Peicheng Wu, Zhenyu Bu, Runze Ma, Lin Du
LensVLM is an inference framework and post‑training recipe that lets Vision‑Language Models (VLMs) process compressed images of text by selectively expanding only the relevant parts back to full resolution. Using Qwen3.5‑9B‑Base, LensVLM achieves accuracy comparable to full‑text models at 4.3× compression and outperforms other compression baselines up to 10.1× across seven text QA benchmarks, while also improving performance on multimodal document and code tasks as compression increases.
By Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan, Christopher Fifty, Jang-Hyun Kim, Xianzhi Du, Zhe Gan, Vivek Rathod, Bhuwan Dhingra
DuplexSpeechBench-Document Grounding (DSB‑DG) is a benchmark that evaluates how well voice agents can ground their responses in external documents across five professional domains. It focuses on three failure modes—Context Saturation, Grounding Decay, and Proactive Grounding—and includes 1,636 adversarially verified QA pairs from 50 documents. The benchmark measures grounding accuracy, hallucination, and response latency, revealing that cascaded pipelines achieve the highest accuracy while open‑weight systems suffer from abrupt capacity collapse and multi‑turn decay.
By Puneet Mathur, Nedim Lipka, Zeyu Jin, Dinesh Manocha
LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.
By Jaeyun Shin, Hangeol Chang, Jong Chul Ye
Vision‑language models (VLMs) are increasingly used as judges to select the best image from a set of generated options, yet their decisions are usually validated only by score agreement with human ratings rather than by the images they return. In this study, the authors audited VLM judges on 300 culturally situated prompts, comparing the returned image with unseen human ratings and with random choice, and repeated each decision with reordered candidates. They found that a 4‑B parameter judge barely outperforms random selection, shows significant position bias, and is highly sensitive to candidate ordering, while an 8‑B judge reduces position bias and outperforms a CLIP similarity baseline but still discards many good decisions when filtered.
By Huichan Seo
MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation.
"whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."
By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen
The paper introduces the Separable Law, a framework that predicts how vision‑language model performance varies with language backbone size and visual token count. By fitting this law to 26 InternVL and QwenVL models across a range of backbone sizes (1B–72B) and image resolutions (224–8K pixels), the authors show that some question types scale predictably with model capacity while others do not. The law also provides a closed‑form rule for allocating compute between backbone size and visual tokens, helping to choose near‑optimal model and image sizes under a fixed budget.
By Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li
Architectural Sampling is a training‑free technique that improves test‑time scaling for frozen vision‑language models by generating diverse candidates through distinct forward computations. It achieves this by reusing selected blocks of decoder layers, varying block location and repetition count, thereby creating computational diversity without updating weights or adding parameters. Experiments on five Qwen checkpoints across twelve multimodal benchmarks show that this method raises pass@9 by an average of 6.58 percentage points over standard temperature sampling, with early‑layer reuse delivering the strongest gains and lower lexical overlap in the generated candidates.
By Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza
The paper introduces Token Communication (TokCom) as a native interface for Collaborative Embodied Artificial Intelligence (CEAI), where tokens act as compact semantic carriers and inference units for generative foundation models. It outlines a TokCom-assisted CEAI framework featuring a task‑adaptive communication protocol that includes a compact codebook, syntax rules, and contextual examples to guide token distillation and reconstruction over wireless channels. A case study on collaborative object transport shows that TokCom significantly reduces source payload bit consumption while maintaining task efficiency and robustness to noisy channels.
By Peng Yi, Ying-Chang Liang
Selection-Based Structured Reasoning (SSR) is a framework that replaces free-form reasoning with the selection of pre-specified natural‑language reasoning candidates, allowing multimodal agents to choose from a set of reusable high‑level reasoning traces. By scoring these candidates in parallel using a shared context KV cache, SSR eliminates the need for an auxiliary task head and reduces inference cost. Experiments on seven multimodal search benchmarks with 2B and 4B models show that SSR maintains competitive success rates while cutting per‑turn reasoning latency by over 90% and overall inference latency by 28–54%.
By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
The paper introduces 3D-Prog, a framework that adapts powerful 2D vision‑language models (VLMs) for reliable 3D understanding, manipulation, and generation. It does so by employing Canonical Coordinate Framing (CCF) to anchor inputs and outputs in a shared Euclidean coordinate system, resolving axis ambiguity, metric scale, and reference issues, and Task‑Adaptive Feedback (TAF) to close the reasoning loop with dynamic, task‑specific feedback. Together, CCF and TAF enable 2D VLMs to perform open‑vocabulary 3D tasks without retraining, yielding consistent, interpretable, and high‑quality results across diverse 3D scenarios.
By Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel
DuoMind is a distributed hierarchical framework that enables multi-robot coordination by combining vision-language-action models for low‑level execution with vision-language models for high‑level reasoning and inter‑robot communication. Each robot’s orchestrator processes task instructions, local observations, and messages from peers to generate precise action commands and semantic messages for other robots. The authors introduce RoboPoly, a benchmark of long‑horizon manipulation tasks, and show through experiments that DuoMind improves multi‑robot task performance, with ablation studies highlighting the roles of hierarchical orchestration and semantic communication.
By Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
The paper introduces IMAVB, a 500‑clip benchmark that tests whether omnimodal large language models can detect when a textual premise contradicts their visual or audio input. Experiments on eight open‑source models and Gemini 3.1 Pro reveal a Representation‑Action Gap: internal states encode mismatches, yet the models rarely reject false premises, exhibiting under‑rejection or over‑rejection. A probe‑guided logit adjustment improves rejection behavior, suggesting the main bottleneck is in translating perception to action rather than in perception itself.
By Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu