arXiv AI By Guangsheng Yu, Litianyi Zhang, Qin Wang, Xu Wang, Mingyuan Li, Shaoxiong Ji, Ren Ping Liu, Massimo Piccardi

Diversity Combining for Multi-Path LLM Reasoning

Read the original on arXiv AI →

The paper treats multi‑path large language model reasoning as a diversity‑combining problem, analogous to noisy channel observations in wireless communications. It shows that the optimal symmetric linear combiner of latent embeddings is uniform, justifying majority vote in standard self‑consistency while allowing for weighting or pruning when prompt‑template branches are heterogeneous. By reducing path correlation through prompt‑template diversity, the authors propose an Adaptive‑K rule that selects an optimal number of reasoning paths, preserving most of the accuracy achieved with a fixed large number of paths across multiple models and benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

The paper introduces Bayesian backward reasoning as a label‑free anchor for multi‑agent collective decision‑making. By constructing reverse posteriors from explicit likelihoods, the authors obtain differently factorized approximations of the underlying posterior, reducing shared errors among agents. Using Jensen‑Shannon divergence to rank agents, they propose three aggregation strategies—hard selection (MinJS), soft reweighting (FwdJS), and log‑linear fusion (LogLin)—which consistently outperform baseline methods on the DDXPlus benchmark across five LLM backbones, especially when agents disagree.

By Ken Chen, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv Machine Learning
Sep 23

Confidence Composition for Multiagent Language Model Systems

The paper addresses the lack of system‑level confidence estimates in multiagent language model systems such as collaborative reasoning and debate. It introduces confidence composition methods, including confidence‑aware routing and log‑odds pooling, to combine agent confidences while maintaining selective utility and probabilistic reliability. Experiments on five benchmarks with diverse model pairs show that gated‑fusion techniques improve AUARC and Brier scores compared to single‑agent and standard debate baselines, and a shared dependence discount further enhances reliability.

By Ali Elahi, Michael J. Curry, Barbara Di Eugenio