arXiv AI

Lexical discovery in unknown environments orchestrated by Large Language Models

arXiv:2607. 22591v1 Announce Type: new Abstract: Populations of autonomous agents deployed in unknown environments (e.

arXiv Computer Vision
Aug 25

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE introduces a visual‑aware sparse Mixture‑of‑Experts framework for embodied referring expression grounding, enabling an agent to navigate real environments and localize a target object from natural language instructions. By routing visual information through specialized experts, the method produces discriminative representations for both navigation views and candidate objects, unlike prior approaches that use a single vision encoder. Experiments on the REVERIE and SOON datasets show that ViSMoE surpasses existing state‑of‑the‑art methods.

By Shuo Feng, Piji Li
arXiv AI
Jun 8

Dual Latent Memory for Visual Multi-agent System

arXiv:2602. 00471v2 Announce Type: replace Abstract: While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs.

By Xinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin, Cheng Yang, Yongbo He, Yihao Hu, Jiangning Zhang, Cheng Tan, Xiaobin Hu, Shuicheng Yan
arXiv AI
Jun 9

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.

By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong
arXiv AI
Oct 2

Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.

By Gabriel Turinici
arXiv Machine Learning
Sep 30

Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

arXiv:2602.15382v3 Announce Type: replace-cross Abstract: Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal st...

By Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, Xiaoqian Wang, Jing Gao