arXiv AI

When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference

arXiv:2606. 08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference.

arXiv AI
Aug 26

Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling

The paper proposes a transparent, user‑configurable rule for selecting arguments in deliberative polls, replacing opaque learned rankers. It formalises argument selection over bipolar justification sets, introduces seven civic recommender criteria, and presents a one‑hop reversed endorsement flow rule that meets them. Experiments on 17,000 simulated runs show the rule performs comparably to random on coverage but outperforms other methods on endorsement mass and robustness under adversarial pressure.

By Muntaser Syed, Markus Zanker, Marius Silaghi
arXiv Computation and Language
Sep 14

Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding

Chopthin-Consensus Power Sampling (CCPS) is a new inference-time decoding method for large language models that uses the Chopthin resampler to preserve diversity among particle trajectories. By enforcing an upper bound on weight ratios instead of equal-weight resampling, CCPS maintains a richer set of distinct reasoning paths and guarantees a lower bound on effective sample size. Coupled with a semantic-majority selection mechanism, CCPS achieves higher oracle coverage and matches or surpasses baseline accuracy on multiple reasoning benchmarks.

By Minoo Ahmadi, Seyedarmin Azizi, Erfan Baghaei Potraghloo, Mehdi Kamal, Massoud Pedram
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu
arXiv AI
2d ago

ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

ReSolve is a training‑free inference method that reuses candidate reasoning by selectively moderating generative outputs. It examines existing derivations when candidates disagree or lack a parseable answer, then incorporates new solutions into a bounded loop. On 130 competition‑mathematics problems, ReSolve achieves 100 and 99 correct answers with significantly fewer tokens than eight‑sample self‑consistency, while a controlled ablation shows that visible derivations improve accuracy.

By Bangji Yang, Jiajun Fan, Hongba Ma, Xi Zhu, Weizhi Zhang, Minghao Guo, Ye Li, Hamid Palangi, Jiaxuan You
arXiv AI
Sep 12

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

The paper introduces Bayesian backward reasoning as a label‑free anchor for multi‑agent collective decision‑making. By constructing reverse posteriors from explicit likelihoods, the authors obtain differently factorized approximations of the underlying posterior, reducing shared errors among agents. Using Jensen‑Shannon divergence to rank agents, they propose three aggregation strategies—hard selection (MinJS), soft reweighting (FwdJS), and log‑linear fusion (LogLin)—which consistently outperform baseline methods on the DDXPlus benchmark across five LLM backbones, especially when agents disagree.

By Ken Chen, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
arXiv AI
Sep 4

Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.

By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason
arXiv AI
Sep 3

Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting

Loom is a generative consensus framework designed for real‑world root‑cause analysis (RCA) that combines open‑form hypotheses from modular heuristics with a lightweight large language model (LLM) synthesis step. It projects hypotheses into a continuous embedding space and uses an iterative centroid‑based reweighting algorithm to resolve conflicts, producing a single consensus that is then synthesized by one LLM call. On the OpenRCA benchmark Loom matches state‑of‑the‑art autonomous agents on some datasets while achieving significantly higher efficiency—about 26× faster and 33× faster with an 8B‑parameter synthesizer. whyItMatters":"Loom demonstrates how embedding‑space reweighting can bridge the gap between statistical rigor and expressive LLMs, enabling efficient, trustworthy RCA in industrial settings."

By Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat
arXiv AI
Aug 14

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608. 12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation.

By Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
Aug 19

A decodability criterion predicts when hidden-state selection beats majority voting in large language models

The paper introduces CASE, a dynamic selection combiner that uses a linear gate trained on answer-token hidden states to choose the best candidate answer from a large language model’s samples. It proposes decodability, a leakage‑free metric that predicts when hidden‑state selection will outperform majority voting, achieving a strong correlation (r=0.75) with accuracy gains. CASE improves accuracy by up to 19 points on medium‑difficulty and 16.8 points on hard questions across general and medical LLMs, and its predictive power transfers to unseen scientific domains.

By Zhixiang wang, Ziliang Hong, Ulas Bagci