arXiv AI

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

arXiv:2608. 13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate.

arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv Machine Learning
Sep 30

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu