No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
arXiv:2608. 13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate.
Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model becomes incorrect under the new one.
arXiv:2608. 13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate.
arXiv:2602. 20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output.
The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.
UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."
The paper introduces Test-Time Calibration Learning (TTCL), a label‑free framework that adapts both reasoning accuracy and verbalized confidence of large language models directly on unlabeled target‑task data. TTCL generates self‑supervision signals from multiple model responses, enabling calibration without ground‑truth labels and proving theoretically as a bounded surrogate for the ideal calibration objective. Experiments on mathematical reasoning and factual question answering show consistent improvements, with base models gaining an average 40.13% in accuracy and 70.80% in ECE reduction across eight benchmarks, and further gains under domain shift.
arXiv:2607. 17531v1 Announce Type: cross Abstract: Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with broadly improving LLM reasoning, yet its gains are uneven and sometimes negative.
arXiv:2606. 19549v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) makes it cheap to train many domain- and task-specific language model adapters, but whether two adapters can be merged is usually discovered only after both have been fully trained and evaluated.
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
Rolling Conformal Prediction (rolling‑CP) is a distribution‑free predictive inference method designed for sequential model training. It calibrates each incoming observation against the current predictor and incorporates it into future training, eliminating the need for data splitting. For exchangeable data, rolling‑CP guarantees marginal coverage with a universal factor‑two bound, and for i.i.d. streams it provides high‑probability training‑conditional validity over time, improving to the target level under stability conditions.
arXiv:2608. 15445v1 Announce Type: new Abstract: When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization.
arXiv:2509. 17314v4 Announce Type: replace-cross Abstract: Software increasingly relies on the emergent capabilities of Large Language Models (LLMs), from natural language understanding to program analysis and generation.
arXiv:2608. 16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail.