arXiv AI By Indranil Halder, Cengiz Pehlevan

Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling

Read the original on arXiv AI →

arXiv:2512. 19905v3 Announce Type: replace-cross Abstract: Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
6d ago

Large Language Bayes Is Not Reparameterisation-Invariant

Large Language Bayes (LLB) samples probabilistic programs from a language model, runs approximate inference on each, and averages them weighted by an exponentiated evidence bound. The authors demonstrate that this weighting is not invariant to reparameterisation, unlike the log marginal likelihood, leading to significant discrepancies in weights across different program formulations. These discrepancies can reach up to 31.9×, affect Bayes factors, and introduce controlled errors in posterior estimates.

By Jian Xu
arXiv AI
Oct 1

Revisiting scaling laws for reward optimization

The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.

By Ali Aouad, Aymane El Gadarri, Vivek F. Farias