Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication
arXiv:2509. 11208v3 Announce Type: replace-cross Abstract: Transformers used for evidence-grounded binary adjudication (e.
arXiv:2507. 11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently.
arXiv:2509. 11208v3 Announce Type: replace-cross Abstract: Transformers used for evidence-grounded binary adjudication (e.
arXiv:2607. 18804v1 Announce Type: new Abstract: In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations.
The paper introduces F-ICL, a benchmark that measures in‑context algorithmic reasoning in language models by exhaustively enumerating 86 million valid programs of length ≤13 on a Turing‑complete machine and computing the exact posterior under a bounded Levin–Solomonoff prior. Unlike typical benchmarks, F‑ICL provides a distributional reference rather than just answers, allowing the evaluation of models’ inductive priors. Across 105 configurations of models ranging from 0.8 B to 675 B parameters, models achieve up to 92 % accuracy, yet many still deviate from the Bayes‑optimal reference, and the study derives theoretical bounds on cumulative loss for predictors with positive prior weight on the reference.
arXiv:2607. 17060v1 Announce Type: new Abstract: A Bayes-filtered transformer (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task.
The paper presents a Bayesian framework that unifies several large‑language‑model training and evaluation paradigms—supervised fine‑tuning (SFT), few‑shot in‑context learning (ICL), and KL‑regularized reinforcement learning (RLHF/RLVR). It shows that each method can be viewed as a two‑step process: first constructing a Bayes or Gibbs posterior over outputs or actions using a prior and a utility signal, then approximating this posterior via a forward‑KL projection onto a parametric family. The authors formalize ICL and SFT as amortized weight projections, and demonstrate that reward‑weighted SFT, reward‑weighted ICL, and advantage‑weighted SFT are all special cases of forward‑KL projection onto reward‑induced posteriors, while also outlining where these equivalences hold and where they break down.
arXiv:2507. 01414v2 Announce Type: replace Abstract: We introduce a new family of toy problems that combine features of linear-regression-style continuous in-context learning (ICL) with discrete associative recall.
The paper investigates feature priming in high‑dimensional online linear regression, showing that estimating feature weights from past data and refitting a minimum‑norm predictor can lead to regret that scales with sparsity rather than ambient dimension. It provides a negative answer to a COLT 2023 open problem by proving that three natural priming rules incur ≥Ω(min{T,√d}) regret against a zero‑loss one‑sparse comparator, due to cheap nuisance interpolation that underweights truly predictive coordinates. The authors also identify conditions under which regret is governed by data rank and present constructions that achieve tight univariate rates, while noting that the multivariate case remains unresolved.
arXiv:2607. 26504v1 Announce Type: new Abstract: Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference.
arXiv:2607. 18422v1 Announce Type: new Abstract: Overparameterized models often have continuous parameter symmetries, so different parameters define the same predictor.
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
arXiv:2603.05335v3 Announce Type: replace-cross Abstract: Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate...
arXiv:2607. 27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive.