ReSolve is a training‑free inference method that reuses candidate reasoning by selectively moderating generative outputs. It examines existing derivations when candidates disagree or lack a parseable answer, then incorporates new solutions into a bounded loop. On 130 competition‑mathematics problems, ReSolve achieves 100 and 99 correct answers with significantly fewer tokens than eight‑sample self‑consistency, while a controlled ablation shows that visible derivations improve accuracy.
By Bangji Yang, Jiajun Fan, Hongba Ma, Xi Zhu, Weizhi Zhang, Minghao Guo, Ye Li, Hamid Palangi, Jiaxuan You
arXiv:2606. 08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference.
By Yasushi Sakai, Allen Song, Kent Larson
arXiv:2608.23086v1 Announce Type: new
Abstract: Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to prioritize h...
By Rounak Sharma, Ananya B. Sai, Soumyabrata Pal
The paper introduces Dual-Seed Comparison (DSC), a protocol that uses two independent LLM-generated seeds to reduce systematic bias in probabilistic sampling. DSC constructs a bit sequence from the character-level ordinal values of the seeds, normalizes it into a pseudo-uniform variate, and maps it to the target distribution via the inverse cumulative distribution function. Empirical results show DSC outperforms existing methods in 96% of evaluated settings and enhances distributional control in tasks like MCQ generation and attribute-constrained text-to-image prompting.
By Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui
The paper introduces a formal framework for Simulation-Augmented Generation (SAGE), a method that simulates individual viewpoints to answer contentious queries more representatively. By applying the metric proportional justified representation+ (mPJR+) axiom from proportional clustering, the authors prove that only a small number of simulations (n ≪ n_H) and dynamic routing to an even smaller subset (k ≪ n) are sufficient to approximate proportional representation for a large population. Empirical results on political and personal advice domains show that their routing algorithm outperforms k‑means and random selection baselines in achieving higher mPJR+ satisfaction rates.
By Sonja Kraiczy, Smitha Milli, Ratip Emin Berker, Avinandan Bose, Brandon Amos, Jamelle Watson-Daniels, Maximilian Nickel, Edith Elkind, Ariel D. Procaccia
arXiv:2608. 14420v1 Announce Type: new Abstract: Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time.
By Haohui Yang, Jiaxing Sun, Xiujun Ma
arXiv:2609.31857v2 Announce Type: replace
Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled....
By Junxuan Li, Arko Mukherjee, Soumyabrata Pal
arXiv:2502. 11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it.
By Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng
arXiv:2511. 12309v2 Announce Type: replace-cross Abstract: Self-consistency (SC) is a widely used test-time inference technique for improving performance in chain-of-thought reasoning.
By Austin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov, Ievgen Redko
Chopthin-Consensus Power Sampling (CCPS) is a new inference-time decoding method for large language models that uses the Chopthin resampler to preserve diversity among particle trajectories. By enforcing an upper bound on weight ratios instead of equal-weight resampling, CCPS maintains a richer set of distinct reasoning paths and guarantees a lower bound on effective sample size. Coupled with a semantic-majority selection mechanism, CCPS achieves higher oracle coverage and matches or surpasses baseline accuracy on multiple reasoning benchmarks.
By Minoo Ahmadi, Seyedarmin Azizi, Erfan Baghaei Potraghloo, Mehdi Kamal, Massoud Pedram
The paper proposes a transparent, user‑configurable rule for selecting arguments in deliberative polls, replacing opaque learned rankers. It formalises argument selection over bipolar justification sets, introduces seven civic recommender criteria, and presents a one‑hop reversed endorsement flow rule that meets them. Experiments on 17,000 simulated runs show the rule performs comparably to random on coverage but outperforms other methods on endorsement mass and robustness under adversarial pressure.
By Muntaser Syed, Markus Zanker, Marius Silaghi
arXiv:2601. 21522v2 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials.
By Sagi Meir, Tommer D. Keidar, Noam Levi, Shlomi Reuveni, Barak Hirshberg