The Capability Manifold and ML Scaling Laws
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insuffi...
The paper introduces a capability manifold, a multidimensional framework that maps downstream capabilities—such as reasoning, retrieval, planning, and adaptation—to pre‑training, post‑training, and test‑time resources via bounded scaling functions. It provides analytical Jacobians to quantify how sensitive each capability is to changes in resources and their interactions. By embedding existing Kaplan‑ and Chinchilla‑type scaling laws and test‑time compute into this manifold, the authors demonstrate that these scaling relationships can be unified as trajectories on a common capability manifold.
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insuffi...
The paper introduces SuperValid, a framework that generates out-of-distribution, capability-aligned validation data by distilling core concepts from benchmarks and expanding them into diverse, knowledge-rich texts. By focusing on capability-level performance rather than benchmark-specific metrics, SuperValid’s loss correlates strongly and stably with downstream benchmark results across a wide range of models, scales, and training data distributions. This training‑free metric can be computed during training, enabling model selection, early stopping, and scaling decisions without the need for benchmark evaluation.
arXiv:2607. 00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget?
The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.
The paper introduces the Capability-Driven Multimodal Scaling Law, a cross-family framework that predicts vision-language model (VLM) benchmark accuracy from a low-dimensional textual capability score extracted via PCA. By training over 150 VLMs on 34 large language models across seven families, the authors demonstrate that the law accurately extrapolates transfer rates from 8B to 72B‑parameter backbones, predicts full training trajectories, and generalizes to unseen model families. The study also reveals actionable insights, such as certain textual benchmarks negatively correlating with multimodal performance and base LLMs outperforming instruction-tuned counterparts as VLM backbones due to higher absorption rates.
arXiv:2512. 16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making.
The paper introduces a formal framework for measuring AI propensities—tendencies of models to exhibit particular behaviours—using a bilogistic formulation that identifies an "ideal band" of success probability. It estimates the limits of this band with task‑agnostic rubrics and applies the method to six families of LLMs, showing how shifts in propensity affect task performance. The study finds that propensity estimates from one benchmark predict behaviour on held‑out tasks and that combining propensity with capability metrics yields stronger predictive power than either alone.
arXiv:2606. 00047v1 Announce Type: cross Abstract: Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used during training.
arXiv:2608. 20061v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost.
arXiv:2608. 13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.
arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.
The paper introduces Switching LoRA Adapters as a Tool (SLAaaT), a method that lets agents dynamically switch between specialized LoRA adapters during a trajectory. By applying this to two synthetic coding tasks, the authors show that agents can solve problems they previously failed, autonomously select strategies that outperform a human heuristic, and reduce the capability tax by up to 18× compared to using a single adapter. SLAaaT also outperforms spawning subagents in both task performance and token efficiency.