arXiv Computation and Language

What Confidence Routing Is Actually Doing: Auditing Routing, Calibration, and Commitment in Multi-Agent Deliberation

arXiv AI
Aug 18

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

arXiv:2608. 14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost.

By Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang, Vikram Vasudevan, Huihuo Zheng, Venkatram Vishwanath, Rajeev Thakur
arXiv AI
1d ago

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

The paper audits whether routing entropy in Attention‑Residual transformer variants (Swin‑Tiny and DeiT‑Small) trained on CIFAR‑10/100 can signal prediction uncertainty beyond model confidence. Three tests examine the presence, consistency, and predictive power of routing signals, while a sensitivity audit measures how much injected effect the probes recover. Results show no significant improvement over confidence alone, with only modest recovery of injected signals and no consistent gains across seeds or metrics.

By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi
arXiv AI
Sep 25

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

MeshHeal is a fully decentralized self‑healing framework for decentralized LLM‑based multi‑agent systems that addresses gray failures—situations where an agent remains responsive but its task‑solving quality degrades. It operates on two timescales: a fast adaptive hierarchy that escalates uncertain or low‑scoring outputs to committee review and correction, and a slow peer‑relative detector that aggregates scores to distinguish persistent degradation from normal variation, triggering mandatory review and eventual exclusion of degraded agents while allowing recovered agents to rejoin. MeshHeal’s evaluation, using Model‑Backed MAS Evaluation, shows it achieves higher degraded‑phase accuracy (0.839) on BBH, MATH, and MMLU‑Pro with fewer tokens per task compared to the baseline Symphony.

By Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou
Hugging Face Trending Papers
Aug 11

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%.

arXiv AI
Jul 20

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

arXiv:2607. 15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones.

By Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur
arXiv Machine Learning
4d ago

Reinforcement Learning of Communication in a Mesh of Small Language Models

The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.

By Mehmet Kerem Turkcan