arXiv AI

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

The paper introduces Bayesian backward reasoning as a label‑free anchor for multi‑agent collective decision‑making. By constructing reverse posteriors from explicit likelihoods, the authors obtain differently factorized approximations of the underlying posterior, reducing shared errors among agents. Using Jensen‑Shannon divergence to rank agents, they propose three aggregation strategies—hard selection (MinJS), soft reweighting (FwdJS), and log‑linear fusion (LogLin)—which consistently outperform baseline methods on the DDXPlus benchmark across five LLM backbones, especially when agents disagree.

arXiv AI
1d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano
arXiv AI
Sep 4

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu
arXiv AI
Aug 5

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

arXiv:2608. 03506v1 Announce Type: new Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace.

By Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho
arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
Hugging Face Trending Papers
Aug 4

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer.

arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv Machine Learning
Sep 23

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

By Ali Elahi, Michael J. Curry, Barbara Di Eugenio