The paper introduces RUPA, a trajectory‑level uncertainty quantification framework for large language model agents. RUPA models an agent’s execution as a directed graph of reasoning states, tool interactions, and environment feedback, then propagates uncertainty across this graph to capture long‑range dependencies. Experiments on benchmarks such as τ‑2, Terminal‑Bench‑2, and GAIA show that RUPA outperforms existing methods, enabling earlier failure detection and more reliable agent execution.
By Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv:2607. 02186v1 Announce Type: new Abstract: Software development is a complex task that demands cooperation among agents with diverse roles.
By Temitayo Olamilekan Ogunsusi, Lijun Qian, Xishuang Dong
arXiv:2607. 07989v1 Announce Type: cross Abstract: Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed structure also introduces new challenges in diagnosing system-level failures.
By Yufei Xia, Anjun Gao, Yueyang Quan, Zhuqing Liu, Minghong Fang
arXiv:2607. 10811v1 Announce Type: cross Abstract: AI engineering is shifting from passive text generation by large language models (LLMs) to agent-driven task execution, creating new reliability challenges for long-horizon tasks under resource constraints and environmental uncertainty.
By Kai Yu, Lu Chen, Hanqi Li
arXiv:2607. 25877v1 Announce Type: new Abstract: This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification.
By Bart Custers, Koorosh Aslansefat
arXiv:2608. 14707v1 Announce Type: new Abstract: As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge.
By John Knowlton, Aritra Guha, Risto Miikkulainen