arXiv:2601. 21754v3 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.
By Haoyu Wang, Guozheng Ma, Shugang Cui, Yilun Kong, Haotian Luo, Li Shen, Mengya Gao, Yichao Wu, Xiaogang Wang, Dacheng Tao
arXiv:2601. 19792v4 Announce Type: replace-cross Abstract: For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical.
By Peter Zeng, Weiling Li, Amie Paige, Zhengxiang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory Zelinsky, Susan Brennan, Owen Rambow
arXiv:2609.00474v1 Announce Type: cross
Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in ma...
By Harini S I, Somesh Singh, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy
ViSMoE introduces a visual‑aware sparse Mixture‑of‑Experts framework for embodied referring expression grounding, enabling an agent to navigate real environments and localize a target object from natural language instructions. By routing visual information through specialized experts, the method produces discriminative representations for both navigation views and candidate objects, unlike prior approaches that use a single vision encoder. Experiments on the REVERIE and SOON datasets show that ViSMoE surpasses existing state‑of‑the‑art methods.
By Shuo Feng, Piji Li
arXiv:2602. 00471v2 Announce Type: replace Abstract: While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs.
By Xinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin, Cheng Yang, Yongbo He, Yihao Hu, Jiangning Zhang, Cheng Tan, Xiaobin Hu, Shuicheng Yan
arXiv:2606. 18271v1 Announce Type: new Abstract: As Earth Observation data generation outpaces downlink bandwidth and human-in-the-loop processing, a widening gap has emerged between onboard collection and actionable ground intelligence.
By Juan Manuel Delfa Victoria, Taran Cyriac John, Andrew W. Herson
arXiv:2601.13247v2 Announce Type: replace-cross
Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural ground...
By Baochang Ren, Yunzhi Yao, Rui Sun, Shuofei Qiao, Ningyu Zhang, Huajun Chen
arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.
By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong
arXiv:2608. 13684v1 Announce Type: new Abstract: This paper describes a neurosymbolic architecture for learning to assemble novel structures using evidence from embodied conversations and task demonstrations.
By Jonghyuk Park, Alex Lascarides, Subramanian Ramamoorthy
The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.
By Gabriel Turinici
arXiv:2602.15382v3 Announce Type: replace-cross
Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...
By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao
arXiv:2607.10744v5 Announce Type: replace
Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models...
By Changfei Fu, Guangcheng Chen, Aoxiang Gu, Haoxiang Liang, Wenjun Xu, Hong Zhang