AI agents

Tool use, function calling, orchestration and the protocols that let models act rather than only answer.

8,935 stories · RSS feed

arXiv AI
Jun 2

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

arXiv:2605. 02122v2 Announce Type: replace-cross Abstract: Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation.

By Akash Bonagiri, Gerard Janno Anderias, Saee Patil, Angelina Lai, Devang Borkar, Gezheng Kang, Ishant Gandhi, Setareh Rafatirad, Houman Homayoun
arXiv AI
Jun 2

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

arXiv:2606. 02359v1 Announce Type: new Abstract: Despite the remarkable progress of Large Language Model (LLM) based Multi-Agent Systems, most research focuses on optimizing coordination topology while largely underexploring the equally critical problem: how to transmit and optimize messages among agents effectively?

By Yao Guan, Lin Wang, Zhihu Lu, Ziyi Wang, Wenzhu Yan, Qiang Duan
arXiv AI
Jun 2

Monitoring Agentic Systems Before They're Reliable

arXiv:2606. 02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape.

By Marisa Ferrara Boston, Glen Hanson, Effi Georgala, JD Hudgens, Heather Frase
arXiv AI
Jun 2

SPARC: Spatial-Aware Path Planning via Attentive Agent Communication

arXiv:2603. 02845v5 Announce Type: replace-cross Abstract: Efficient communication is critical for decentralized Multi-Robot Path Planning (MRPP), yet existing learned communication methods treat all neighboring robots equally regardless of their spatial proximity, leading to diluted attention in congested regions where coordination matters most.

By Sayang Mu, Xiangyu Wu, Bo An
arXiv AI
Jun 2

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

arXiv:2512. 07436v3 Announce Type: replace Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources.

By Hang He, Chuhuai Yue, Chengqi Dong, Mingxue Tian, Hao Chen, Zhenfeng Liu, Jiajun Chai, Xiaohan Wang, Yufei Zhang, Qun Liao, Guojun Yin, Wei Lin, Chengcheng Wan, Haiying Sun, Ting Su