arXiv Machine Learning

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv AI
1d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano
arXiv Machine Learning
Sep 24

Reliable Fusion of Conflicting Experts

The paper introduces a probabilistic‑circuit framework for fusing opinions from multiple black‑box experts in noisy, conflict‑prone environments. It dynamically assigns context‑specific credibility to each expert, allowing reliable aggregation without needing access to their internal models or retraining. Experiments on multiple‑choice question answering with large language models show that this method outperforms individual models and static ensemble baselines, consistently improving predictive accuracy and decision reliability under disagreement.

By Pranuthi Tenali, Sahil Sidheekh, Saurabh Mathur, Vijayalakshmi Saravanan, Erik Blasch, Kristian Kersting, Sriraam Natarajan
arXiv AI
Sep 2

DualStake: Dual-Path Confidence Calibration in Deep Research Agents

DualStake introduces a dual-path confidence calibration for deep research agents, adding step confidence elicitation after each retrieval step. The method shows that evidence confidence (E-Conf) after the final retrieval provides a stronger uncertainty signal than answer confidence (A-Conf), and that A-Conf is largely influenced by E-Conf. By applying margin‑clipped, confidence‑dependent stake rewards, DualStake aligns both E-Conf and A-Conf with answer correctness, improving calibration across multiple QA benchmarks without harming accuracy.

By Yinuo Xu, Yuwei Liang, Jianjie Cheng, Meng Wang, Yongcan Yu, Shuo Lu, Jian Liang
arXiv AI
Sep 4

Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

The paper introduces R$^2$-MAD, a framework that enhances multi-agent debate by giving agents an experience memory from past debates. It uses a debate-state-aware retrieval policy to adjust concept priors based on current consensus, and derives confidence weights from retrieved experiences to modulate peer influence. Experiments demonstrate consistent improvements over existing single-agent and MAD baselines.

By Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang
arXiv AI
Aug 24

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

DirEAG introduces a Dirichlet Evidence Aggregation technique to calibrate verbalized confidence in large language models performing mathematical reasoning. By converting each elicited answer-confidence pair into calibrated soft evidence over candidate answers and a null state, it addresses prompt- and task-dependent bias that simple averaging or heuristic aggregation cannot handle. Experiments on GSM8K, SVAMP, and GSM-Hard with Qwen, Mistral, and Gemma models demonstrate that DirEAG achieves better calibration while maintaining competitive answer selection compared to existing methods.

By Haorui Xu, Yuzhou Zhu, Liyuan Gao
arXiv AI
Sep 4

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason