The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.
By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv:2610.08216v1 Announce Type: new
Abstract: Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pool...
By Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu
arXiv:2610.06977v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-sou...
By Xiaojun Jia, Simeng Qin, Yiming Li, Jie Liao, Sensen Gao, Ke Ma, Yang Liu, Xiaochun Cao
arXiv:2610.07339v1 Announce Type: cross
Abstract: Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidan...
By Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran
arXiv:2610.07460v1 Announce Type: cross
Abstract: Inserting objects into existing 3D scenes requires more than selecting a plausible location:
the inserted object must also fit local geometry while...
By Tzu-Hsin Hsieh, Ricardo Marroquim
arXiv:2610.07654v1 Announce Type: cross
Abstract: On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Rece...
By Jian Luo, Kehan Qi, Qingqiao Hu, Meilong Xu, Jiacheng Qiu, Weimin Lyu, Jiawei Zhou, Chao Chen
arXiv:2610.07758v1 Announce Type: cross
Abstract: Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at...
By Hyeongheon Cha, Hyungjun Yoon, Sung-Ju Lee
arXiv:2610.08183v1 Announce Type: cross
Abstract: Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performa...
By Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang
arXiv:2610.08331v1 Announce Type: cross
Abstract: The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in ex...
By Heyam Bin Jahlan Areej Alhothali Abeer Alhothali
arXiv:2610.08482v1 Announce Type: cross
Abstract: Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited,...
By Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, Xiaojuan Li, Mingrui Yang
arXiv:2610.08659v1 Announce Type: cross
Abstract: Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-bas...
By Suxin Ji, Hungtao Wan, Mingjun Liu, An Zhang
arXiv:2610.04206v2 Announce Type: replace
Abstract: Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Archit...
By Uttamasha Monjoree, Wei Yan
arXiv:2510.17826v2 Announce Type: replace-cross
Abstract: Building a working mental model of a protein typically requires weeks of reading, cross-referencing crystal and predicted structures, and ins...
By Carles Navarro, Mariona Torrens, Philipp Th\"olke, Stefan Doerr, Gianni De Fabritiis
arXiv:2607.02092v4 Announce Type: replace-cross
Abstract: Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves furt...
By Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Ningwei Bai, Qichen Yin, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Yi Dong, Zezhi Tang
arXiv:2605.23328v3 Announce Type: replace
Abstract: Emotion Recognition in Conversation is a core component of affective computing, while current sign language emotion datasets primarily focus on iso...
By Yusong Wang, Keyu Mao, Takao Obi, Minghao Shao, Kotaro Funakoshi
arXiv:2610.07647v1 Announce Type: cross
Abstract: Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investiga...
By Seymanur Akti, Alexander Waibel
arXiv:2610.07545v1 Announce Type: new
Abstract: We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retra...
By Zubair Ajmal Farooq, Diptesh Kanojia
arXiv:2610.07466v1 Announce Type: new
Abstract: Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using p...
By Payal Mohapatra, Haodong Yang, Yueyuan Sui, Stephen Xia, Benjamin Lundell, Qi Zhu
arXiv:2610.07269v1 Announce Type: cross
Abstract: Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles thr...
By Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah
arXiv:2610.07572v1 Announce Type: cross
Abstract: In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every q...
By Xi Ding, Naichen Shi, Jiawei Zhang