arXiv AI

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

arXiv AI
Aug 24

Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning

The paper introduces a method to improve test-time scaling (TTS) for large language models by using multi-agent systems (MAS) to split long reasoning chains into manageable contexts. A new dataset, M500, containing 500 multi-agent collaborative reasoning traces, is used to fine‑tune open‑source models, enabling them to learn collaborative patterns and outperform their base versions. An adaptive scaling strategy with a "CEO" agent is proposed to dynamically guide reasoning depth, and experiments in the AgentVerse framework confirm the effectiveness of the approach.

By Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N. Metaxas, Tong Che
arXiv AI
Sep 12

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

The paper introduces Bayesian backward reasoning as a label‑free anchor for multi‑agent collective decision‑making. By constructing reverse posteriors from explicit likelihoods, the authors obtain differently factorized approximations of the underlying posterior, reducing shared errors among agents. Using Jensen‑Shannon divergence to rank agents, they propose three aggregation strategies—hard selection (MinJS), soft reweighting (FwdJS), and log‑linear fusion (LogLin)—which consistently outperform baseline methods on the DDXPlus benchmark across five LLM backbones, especially when agents disagree.

By Ken Chen, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
arXiv AI
1d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano