Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces Adaptive Self-Consistency (ASC), a method that treats large language models as grey-boxes by leveraging the full answer distribution from log‑probabilities rather than single sampled answers. It formulates inference as sequential mode identification with distribution‑valued observations, proving that this approach never performs worse than traditional black‑box sampling and can stop earlier. The proposed ASC‑D algorithm, a betting‑based stopping rule, achieves significant reductions in required trajectories—up to 95.6% fewer—while attaining the best fixed‑budget correct‑certification rates on MMLU‑Redux across three open‑source models.
arXiv:2602. 05395v2 Announce Type: replace-cross Abstract: A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached.
arXiv:2606. 17832v1 Announce Type: new Abstract: Large language models (LLMs) are often hypothesized to perform implicit Bayesian inference, yet a key coherence condition, the martingale property of predictive beliefs, has been shown to fail in controlled synthetic in-context learning settings.
Flow Reasoning Models (FRMs) are a new framework that turns continuous flow models into efficient recurrent reasoners for structured tasks. By self‑conditioning a flow model on its own past outputs, FRMs iteratively refine solutions, allowing parallel decision making and revision. The authors introduce Fixed‑Point Forcing (FPF) to mitigate exposure bias at deeper recursion, and report near‑perfect solve rates on Sudoku‑Extreme, Zebra, and Maze‑Unique, outperforming existing masked‑diffusion and specialized baselines while using far fewer inference FLOPs.
arXiv:2606.16011v2 Announce Type: replace Abstract: Standard accuracy benchmarks evaluate whether large language models (LLMs) reach correct answers. However, they do not test whether models maintain...
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.