Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insuffi...
The paper introduces SuperValid, a framework that generates out-of-distribution, capability-aligned validation data by distilling core concepts from benchmarks and expanding them into diverse, knowledge-rich texts. By focusing on capability-level performance rather than benchmark-specific metrics, SuperValid’s loss correlates strongly and stably with downstream benchmark results across a wide range of models, scales, and training data distributions. This training‑free metric can be computed during training, enabling model selection, early stopping, and scaling decisions without the need for benchmark evaluation.
By Quanen Sun, Changxin Tian, Ke Shi, Cai Chen, Cunyin Peng, Jia Liu, Kunlong Chen, Zhiqiang Zhang, Jun Zhou
arXiv:2607. 00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget?
By Alex Fogelson, Zachary A. Brown, Hans Gundlach, Jayson Lynch, Neil Thompson
The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.
By Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava
The paper introduces the Capability-Driven Multimodal Scaling Law, a cross-family framework that predicts vision-language model (VLM) benchmark accuracy from a low-dimensional textual capability score extracted via PCA. By training over 150 VLMs on 34 large language models across seven families, the authors demonstrate that the law accurately extrapolates transfer rates from 8B to 72B‑parameter backbones, predicts full training trajectories, and generalizes to unseen model families. The study also reveals actionable insights, such as certain textual benchmarks negatively correlating with multimodal performance and base LLMs outperforming instruction-tuned counterparts as VLM backbones due to higher absorption rates.
By Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
arXiv:2512. 16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making.
By Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava