arXiv AI By Jingkai Huang, Yunfan Zhang, Will Ma, Weihua Zhou, Zhengyuan Zhou

Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback

Read the original on arXiv AI →

The paper introduces Adaptive Self-Consistency (ASC), a method that treats large language models as grey-boxes by leveraging the full answer distribution from log‑probabilities rather than single sampled answers. It formulates inference as sequential mode identification with distribution‑valued observations, proving that this approach never performs worse than traditional black‑box sampling and can stop earlier. The proposed ASC‑D algorithm, a betting‑based stopping rule, achieves significant reductions in required trajectories—up to 95.6% fewer—while attaining the best fixed‑budget correct‑certification rates on MMLU‑Redux across three open‑source models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 17

From Drift to Coherence: Stabilizing Beliefs in LLMs

arXiv:2606. 17832v1 Announce Type: new Abstract: Large language models (LLMs) are often hypothesized to perform implicit Bayesian inference, yet a key coherence condition, the martingale property of predictive beliefs, has been shown to fail in controlled synthetic in-context learning settings.

By SongEun Kim, Seungyoo Lee, Edwin Fong, Hyungi Lee, Juho Lee
arXiv AI
Sep 2

Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners

Flow Reasoning Models (FRMs) are a new framework that turns continuous flow models into efficient recurrent reasoners for structured tasks. By self‑conditioning a flow model on its own past outputs, FRMs iteratively refine solutions, allowing parallel decision making and revision. The authors introduce Fixed‑Point Forcing (FPF) to mitigate exposure bias at deeper recursion, and report near‑perfect solve rates on Sudoku‑Extreme, Zebra, and Maze‑Unique, outperforming existing masked‑diffusion and specialized baselines while using far fewer inference FLOPs.

By Alec Helbling, Andrey Bryutkin, Mauro Martino, Duen Horng Chau, Nima Dehmamy, Hendrik Strobelt
arXiv AI
Aug 26

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Selective Regenerative Decoding (SRD) is a new inference-time decoding method that improves large language model reasoning by allowing segment-level intervention on candidate trajectories. Instead of discarding or keeping entire trajectories, SRD selectively refines only the degraded suffix while preserving useful prefixes, leading to higher expected trajectory quality and better sample efficiency. Experiments on MATH500, GPQA Diamond, HotpotQA, and AlpacaEval show that SRD matches Best-of-N accuracy with fewer generated tokens and outperforms speculative rejection in low‑compute settings.

By Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung, Yi-an Lai, Arshit Gupta