arXiv Machine Learning

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

arXiv:2605. 22949v3 Announce Type: replace Abstract: Foundation-model pools are increasingly used as black-box responders in coordinated systems where a coordinator must decide which response to trust.

arXiv AI
Sep 3

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.

By Shubhra Mittal
arXiv AI
Sep 18

Diagnose, Recover, Certify: Task Readiness under Hidden Dynamics Changes

The paper introduces a framework for diagnosing and recovering from hidden dynamics changes in deployed control policies, focusing on the problem of task readiness under dormant dynamics drift. It proposes an intervention-based Bayesian method called Evidence‑Gated Matched‑Pulse Transport that localizes faults and estimates actuator effectiveness, enabling agents to certify readiness for future tasks with limited, task‑agnostic interactions. The approach is evaluated on diverse benchmarks, measuring readiness coverage, selective risk, interaction cost, and return, and identifies regimes where transported evidence is decisive.

By Nguyen Viet Tuan Kiet, Huynh Thi Thanh Binh
arXiv AI
Sep 17

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

The paper introduces XConf, an experiential confidence estimator that augments a language model’s current inference with a record of its past graded episodes. By recalling similar past tasks and reflecting on past outcomes, XConf generates confidence scores without accessing logits or updating weights, achieving superior discrimination and calibration across diverse benchmarks. The method demonstrates significant gains in selective prediction, improving success rates on agent tasks by up to 8.7 points.

By Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, Nigel Collier
arXiv Machine Learning
Jun 25

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

arXiv:2606. 25760v1 Announce Type: new Abstract: Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions.

By Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Sep 2

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

The paper introduces an action‑class diagnostic framework for multi‑turn tool‑calling in large language model agents, breaking failures into action‑class miscalibration and action‑execution failure across a four‑class action space (TOOL_CALL, ASK, REFUSE, CONFIRM). It defines a self‑revealing upper bound (Acc GAR) to expose state‑grader masking of miscalibration and shows that miscalibration is a significant, previously hidden failure mode, especially for heavily tool‑trained families. The study demonstrates that calibration can be reshaped by context‑only perturbations, but the effects vary widely across models and perturbation mechanisms, underscoring the need for diagnostics beyond aggregate accuracy.

By Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin