arXiv AI

Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction

arXiv AI
Sep 4

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason
arXiv Machine Learning
Sep 23

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

By Ali Elahi, Michael J. Curry, Barbara Di Eugenio
arXiv Computation and Language
Aug 27

Localize-Then-Decide Guarantees for LLM Judgments

Large language models (LLMs) are increasingly used to evaluate output quality, but guaranteeing agreement with human judgments is difficult. The paper introduces a Localize-Then-Decide framework that first uses conformal prediction to narrow down a shortlist likely to contain the human-preferred response, then applies a calibrated confidence rule to select a single response or abstain. Experiments show this two-stage approach consistently yields higher guarantee success rates and greater coverage than single-stage baselines across various candidate sizes and datasets.

By Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.