arXiv AI

Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling

arXiv:2512. 19905v3 Announce Type: replace-cross Abstract: Recent developments in large language models have shown advantages in reallocating a notable share of computational resource from training time to inference time.

arXiv Machine Learning
6d ago

Large Language Bayes Is Not Reparameterisation-Invariant

Large Language Bayes (LLB) samples probabilistic programs from a language model, runs approximate inference on each, and averages them weighted by an exponentiated evidence bound. The authors demonstrate that this weighting is not invariant to reparameterisation, unlike the log marginal likelihood, leading to significant discrepancies in weights across different program formulations. These discrepancies can reach up to 31.9×, affect Bayes factors, and introduce controlled errors in posterior estimates.

By Jian Xu
arXiv AI
Oct 1

Revisiting scaling laws for reward optimization

The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.

By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv AI
Jul 28

Hierarchical Grading in Large Language Models

arXiv:2607. 22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective.

By T. Shaska
Hugging Face Trending Papers
4d ago

Dynamic Minimax Regret Optimization for Robust LLM Post-Training

The paper introduces DUCB-OGD, an algorithm that couples a Discounted Upper‑Confidence‑Bound sampler with Online Gradient Descent to address dynamic minimax regret in robust large‑language‑model post‑training. It operates under instantaneous mini‑batch‑only bandit feedback, tracking worst‑source performance without re‑evaluating historical data. Experiments on fine‑tuning, preference optimization, and reinforcement learning demonstrate that DUCB‑OGD improves worst‑group robustness with negligible computational overhead.