Hugging Face Trending Papers

No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

Read the original on Hugging Face Trending Papers →

Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model becomes incorrect under the new one.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 30

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv AI
3d ago

Test-time Calibration Learning for Large Language Model Reasoning

The paper introduces Test-Time Calibration Learning (TTCL), a label‑free framework that adapts both reasoning accuracy and verbalized confidence of large language models directly on unlabeled target‑task data. TTCL generates self‑supervision signals from multiple model responses, enabling calibration without ground‑truth labels and proving theoretically as a bounded surrogate for the ideal calibration objective. Experiments on mathematical reasoning and factual question answering show consistent improvements, with base models gaining an average 40.13% in accuracy and 70.80% in ECE reduction across eight benchmarks, and further gains under domain shift.

By Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han