The paper introduces TacEx, a tactile‑curiosity framework that guides reinforcement learning agents to explore contact dynamics by focusing epistemic uncertainty on the tactile channel. By anchoring curiosity to touch, robots learn to manipulate and grasp objects without task rewards or demonstrations, generating an interaction‑dense dataset that supports offline pick‑and‑place policy learning. TacEx also enhances vision‑language‑action models through post‑training, improving downstream performance while remaining sample‑efficient.
By Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza
arXiv:2609.38298v1 Announce Type: new
Abstract: Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial...
By Binchi Zhang, Atrisha Sarkar, Apurva Narayan
arXiv:2609.40041v1 Announce Type: new
Abstract: We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGh...
By Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson, Kishankumar Bhimani, Jochebed Afua Basil
Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context...
Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typica...
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tr...
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relyi...
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities inde...
This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the informati...
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collaps...
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and pr...
Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to...
CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase...
Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio com...
CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.
By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
By Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
PARSEE-VAD is a training‑free online video anomaly detection framework that separates semantic evidence acquisition from score‑state evolution. It uses Proposition‑Aware Reasoning to extract structured propositional evidence from the current causal window and selectively activates more specific queries, while Streaming Evidence Escalation maps this evidence into a compact score‑domain event state and propagates only the bounded state to maintain temporal continuity. Experiments on four benchmarks show strong performance with reduced specialist computation and sparse score‑state propagation, supporting a current‑first principle for streaming multimodal inference.
By Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
RawVLA introduces a streaming neural image signal processor that adaptively renders RAW observations for vision‑language‑action (VLA) policies, focusing on imaging factors that influence embodied behavior. The authors systematically analyze how five ISP dimensions—gain, sensor noise, chromatic response, tonal response, and bit depth—affect action prediction and manipulation success, showing that RAW‑to‑RGB processing significantly shapes outcomes. They also present RawVLA‑Bench, a RAW‑domain manipulation benchmark that evaluates image processing as an explicit variable across clean and adverse conditions, demonstrating that RawVLA maintains performance under standard settings while markedly improving robustness under degraded imaging.
By Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui
The study investigates whether a prosody‑trained representation can improve automatic speech recognition beyond the effect of a trainable fusion mechanism. Using a frozen HuBERT backbone and a 64‑dimensional prosodic representation, the authors compare a baseline recognizer, a fusion model with no auxiliary input, and a fusion model with the learned representation across Buckeye, Switchboard, and AMI IHM datasets. While the fusion model without auxiliary input reduces WER relative to the baseline, adding the prosodic representation yields no significant WER improvement, though the model still depends on the representation for optimal performance.
By Ki Woong Moon, Daniel Brenner
ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.
By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang