arXiv Machine Learning

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

arXiv:2608. 07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated.

arXiv Machine Learning
Sep 21

Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.

By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach
arXiv AI
Sep 2

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.

By Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
arXiv Computation and Language
Sep 4

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.

By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
arXiv AI
Sep 3

Do Large Language Models Capture the Diversity in their Training Data?

The paper investigates whether large language models (LLMs) capture the full diversity of outputs present in their training data. Using an information‑theoretic approach, the authors compare the conditional entropy of model‑generated outputs with that of the training data, finding that LLMs consistently produce outputs with lower conditional entropy across various models, scales, and decoding strategies. They also extend the analysis to image and text‑conditioned generators, propose a post‑hoc correction method based on matrix‑entropy projection to increase conditional diversity, and provide theoretical guarantees and an efficient algorithm for this correction.

By Youqi Wu, Farzan Farnia