Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,396 stories · RSS feed

arXiv Machine Learning
Jul 1

EgoCogNav: Cognition-aware Human Egocentric Navigation

arXiv:2511. 17581v3 Announce Type: replace Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding.

By Zhiwen Qiu, Ziang Liu, Wenqian Niu, Tapomayukh Bhattacharjee, Saleh Kalantari
arXiv AI
Jul 1

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

arXiv:2606. 31329v1 Announce Type: cross Abstract: Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation.

By Dongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung Kim, Jaegul Choo
arXiv AI
Jul 1

Agentic AI Enhances Physician Trust in Clinical Decision Making

arXiv:2606. 30658v1 Announce Type: cross Abstract: Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users.

By Zhiling Yan, Zhe Fang, David J King, Ann Pongsakul, Eashan Adhikarla, Hui Ren, Sunyang Fu, Quanzheng Li, Lifang He, Xiang Li, Hongfang Liu, Yonghui Wu, Lichao Sun