Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computer Vision
6d ago

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

Lang3DSeg introduces a point‑transformer backbone for open‑vocabulary, annotation‑free 3D LiDAR segmentation, trained from scratch without geometric pre‑training. It tackles noise from 2D‑to‑3D label projections by applying a class‑priority rule and truncating projected instances at depth gaps, thereby correcting depth‑ambiguity errors. The method achieves state‑of‑the‑art results on nuScenes (52.8 % mIoU) and SemanticKITTI (41.4 % mIoU) while operating in real‑time on a single LiDAR sweep.

By Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pes\'e, Bing Li
arXiv Computer Vision
6d ago

A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform

The paper reviews end‑to‑end autonomous driving (E2E‑AD) training, framing it as a Data‑Strategy‑Platform system. It surveys recent advances in data pipelines, learning paradigms, and training infrastructures, and discusses how these layers interact to influence model performance, robustness, and deployability. The authors highlight current limitations and propose a future vision that prioritizes data value, foundation‑driven generalization, and integrated training‑testing loops for more robust, scalable, and trustworthy autonomous driving systems.

By Chengkai Xu, Yiming Cui, Jiaqi Liu, Yicheng Guo, Cheng Qin, Geyuan Zhang, Xinwei Dong, Shiyu Fang, Peng Hang, Jian Sun
arXiv Computer Vision
6d ago

Beyond Localization: A Comprehensive Benchmark of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

The paper introduces PCSR-Bench, a benchmark of 84,373 question‑answer pairs derived from 2,600 omnidirectional images across 26 indoor environments, designed to evaluate perspective‑conditioned spatial reasoning (PCSR) in multimodal large language models (MLLMs). It reports a significant perception–reasoning gap, with accuracy dropping from 57.59% on limited field‑of‑view reasoning to as low as 0.64% on open‑ended compositional directional chains. An RL‑based diagnostic study on a 7B‑scale model shows that reward shaping can improve performance to 60.06% on a controlled task, indicating partial plasticity of PCSR capabilities.

By Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras, Xu Zheng
arXiv Computer Vision
6d ago

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.

By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv AI
6d ago

AiSearch: Interactive Multi-Modal Search with VLMs

AiSearch is a flexible multimodal retrieval framework that uses Vision Language Models (VLMs) to enable natural language search over images and videos. It supports interactive search refinement through user feedback, allowing results to be tailored to the user's intent in real time. The system also provides visual benchmarking across multiple VLMs, enabling users to choose the most suitable model for their specific task.

By Ali Koksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck Fong
arXiv AI
6d ago

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

The paper introduces a method to purify LoRA-tuned large language models (LLMs) against backdoor attacks without relying on trigger knowledge, clean references, or retraining. By extracting high‑fidelity backdoor directions and projecting LoRA updates onto orthogonal null spaces in input and output channels, the approach reduces attack success rates from nearly 100% to under 10%. Experiments demonstrate that this null‑space projection preserves both the base model’s general capabilities and the new downstream skills learned through the adapter across various tasks.

By Jianwei Li, Jung-Eun Kim
arXiv Machine Learning
6d ago

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

arXiv:2609.40245v2 Announce Type: cross Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-...

By Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
arXiv AI
6d ago

Supervising Sound Localization by In-the-wild Egomotion

The paper introduces a method for learning binaural sound localization by using egomotion as a supervisory signal. By tracking how a camera’s direction changes relative to a sound source during a video, the authors train an audio model to predict sound directions that align with visual estimates of camera motion derived from multi‑view geometry. They evaluate this approach on a newly proposed dataset of real‑world audio‑visual videos with egomotion, demonstrating that the model can learn from real data and perform well on sound localization tasks.

By Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens