How Inference Compute Shapes Frontier LLM Evaluation
arXiv:2606. 17930v1 Announce Type: new Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving.
arXiv:2607. 00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget?
arXiv:2606. 17930v1 Announce Type: new Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving.
The paper introduces Thomson, a frontier AI model developed through continual learning on open-weight models, aiming to democratize access to high-performance AI. It argues that institutions with limited resources can achieve frontier-level performance by applying a modern mid- & post-training stack, preserving model plasticity and stability while minimizing high-impact interventions. Thomson demonstrates competitive performance across agentic tasks, safety, legal, tax, multilingualism, and large-scale deep research, exhibiting a distinctive π-shaped improvement pattern and effectively mitigating the forgetting problem seen in narrow domain adaptation.
arXiv:2606. 00047v1 Announce Type: cross Abstract: Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used during training.
The paper introduces a capability manifold, a multidimensional framework that maps downstream capabilities—such as reasoning, retrieval, planning, and adaptation—to pre‑training, post‑training, and test‑time resources via bounded scaling functions. It provides analytical Jacobians to quantify how sensitive each capability is to changes in resources and their interactions. By embedding existing Kaplan‑ and Chinchilla‑type scaling laws and test‑time compute into this manifold, the authors demonstrate that these scaling relationships can be unified as trajectories on a common capability manifold.
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insuffi...
arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.
arXiv:2608. 13272v1 Announce Type: new Abstract: A small number of firms based in two states produce the most capable frontier AI models.
arXiv:2601.05280v4 Announce Type: replace-cross Abstract: On the one hand, the question of whether Large Language Models (LLMs) are Solomonoff induction estimators has become an explicit question at...
arXiv:2503. 14499v4 Announce Type: replace Abstract: Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear.
arXiv:2607. 25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data.
The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.
The paper introduces a formal framework for measuring AI propensities—tendencies of models to exhibit particular behaviours—using a bilogistic formulation that identifies an "ideal band" of success probability. It estimates the limits of this band with task‑agnostic rubrics and applies the method to six families of LLMs, showing how shifts in propensity affect task performance. The study finds that propensity estimates from one benchmark predict behaviour on held‑out tasks and that combining propensity with capability metrics yields stronger predictive power than either alone.