arXiv Machine Learning

The Challenger: When Do New Data Sources Justify Switching Machine Learning Models?

arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.

arXiv AI
Jun 10

A Theory of Training Profit-Optimal LLMs

arXiv:2605. 16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure.

By Sophie Hao, William Merrill
arXiv AI
3d ago

Revisiting scaling laws for reward optimization

The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.

By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv Machine Learning
Sep 24

Rolling Conformal Prediction in Sequential Model Training

Rolling Conformal Prediction (rolling‑CP) is a distribution‑free predictive inference method designed for sequential model training. It calibrates each incoming observation against the current predictor and incorporates it into future training, eliminating the need for data splitting. For exchangeable data, rolling‑CP guarantees marginal coverage with a universal factor‑two bound, and for i.i.d. streams it provides high‑probability training‑conditional validity over time, improving to the target level under stability conditions.

By Chen Cheng, Ruiting Liang, Rina Foygel Barber
arXiv Machine Learning
Aug 4

T-TAMER: Provably Taming Trade-offs in ML Serving

arXiv:2509. 22992v2 Announce Type: replace Abstract: As machine learning models continue to grow in size and complexity, efficient serving faces increasingly broad trade-offs spanning accuracy, latency, resource usage, and other objectives.

By Yuanyuan Yang, Ruimin Zhang, Jamie Morgenstern, Haifeng Xu