Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,303 stories · RSS feed

arXiv AI
Jul 9

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

arXiv:2607. 07370v1 Announce Type: cross Abstract: In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution.

By Xufeng Zhao, Fuzhi Yang, Jianhui Chen, Li Gao, Zhang Meng, Jie Gao, Yao Zheng, Wenyu Liu, Menglin Yang, Minqi Gu, Yaru Zhao, Honglin Han, Shihui Su, Zixiao Tang, Liu Liu, Mu Xu, Yang Cai, Wenbin Tang
arXiv AI
Jul 9

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

arXiv:2607. 07375v1 Announce Type: cross Abstract: Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems.

By Chethan Krishnamurthy Ramanaik, Tobias Callies, Michael Hecht, Eirini Ntoutsi
arXiv AI
Jul 9

When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs

arXiv:2607. 07395v1 Announce Type: cross Abstract: Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning improves zero-shot accuracy but often degrades calibration due to entropy-driven overconfidence.

By Tanay Sodha, Aditya Sharma, Ramya Hebbalaguppe, Vinti Agarwal, Pranav Murthy Yeluripaty
arXiv AI
Jul 9

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

arXiv:2607. 07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging.

By Ethan Chung, Chuanjun Zheng, Jasper Tan, Jingxi Li, Haopeng Zhang, Huaijin Chen
arXiv AI
Jul 9

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.

By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez