arXiv:2609.36352v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multip...
By Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng, Haibo Ding, Jun Huan
arXiv:2609.36798v1 Announce Type: cross
Abstract: Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training para...
By Yueran Ma, Ronghao Lin
arXiv:2609.37889v1 Announce Type: cross
Abstract: Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks whil...
By Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan, Da-Wei Zhou
The paper introduces Alignment‑Guided Flow Transformer (AGFT), a framework for Vision‑Language‑Action (VLA) models that explicitly enforces tri‑modal alignment among vision, language, and action through a dedicated alignment loss. AGFT bridges representational gaps across modalities, improving task adaptation and robustness, and employs a flow‑matching objective to reduce inference steps compared to diffusion‑based policies. Experiments on a large benchmark demonstrate that AGFT achieves higher success rates and lower inference latency than state‑of‑the‑art baselines, highlighting tri‑modal alignment as crucial for scalable VLA manipulation.
By Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
arXiv:2501.12632v3 Announce Type: replace-cross
Abstract: Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, wi...
By Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre, Marco Pedersoli, Eric Granger
The paper introduces AttWarp, a lightweight technique that uses a multimodal large language model’s cross‑modal attention to perform rectilinear warping of input images at test time. By reallocating spatial resolution toward query‑relevant regions without altering model weights or architecture, AttWarp preserves global context while making small objects and subtle relationships easier for the model to read. Experiments on five benchmarks and four MLLMs show consistent accuracy gains, improved compositional reasoning, and reduced hallucinations compared to baseline image‑manipulation methods.
By Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim, Madhav Kanda, Hyeonjeong Ha, Svetlana Lazebnik, Heng Ji, Unnat Jain
OVIG is an optimistic verification framework that audits AI training by replaying the process and comparing gradient differences against an empirically calibrated boundary. It treats any gradient difference exceeding this boundary as a malicious deviation. By partitioning training into stride‑s intervals and storing evidence only at interval endpoints, OVIG dramatically reduces off‑chain storage and transmission costs while maintaining zero attack success rate across language, vision, and diffusion workloads.
By Hongxu Su, Jianzhu Yao, Huan Zhang, Xuechao Wang, Pramod Viswanath
arXiv:2609.34911v2 Announce Type: replace-cross
Abstract: Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replann...
By Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye
arXiv:2609.35232v2 Announce Type: replace-cross
Abstract: Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pr...
By Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
arXiv:2609.37638v1 Announce Type: new
Abstract: Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the in...
By Van Bach Nguyen, J\"org Schl\"otterer, Christin Seifer
arXiv:2609.37098v1 Announce Type: cross
Abstract: Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providin...
By Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, Bin Ran
arXiv:2609.37786v1 Announce Type: new
Abstract: Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These represen...
By R\'emi Kazmierczak, Johanne Cohen, Marianne Clausel
arXiv:2603.12717v2 Announce Type: replace-cross
Abstract: Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are de...
By Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar
As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanis...
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introd...
Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated...
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component wi...
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing...
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal...