arXiv AI By Syed Ali Raza Zaidi, Maryam Hafeez

The Capability Manifold and ML Scaling Laws

Read the original on arXiv AI →

The paper introduces a capability manifold, a multidimensional framework that maps downstream capabilities—such as reasoning, retrieval, planning, and adaptation—to pre‑training, post‑training, and test‑time resources via bounded scaling functions. It provides analytical Jacobians to quantify how sensitive each capability is to changes in resources and their interactions. By embedding existing Kaplan‑ and Chinchilla‑type scaling laws and test‑time compute into this manifold, the authors demonstrate that these scaling relationships can be unified as trajectories on a common capability manifold.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 4

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

The paper introduces SuperValid, a framework that generates out-of-distribution, capability-aligned validation data by distilling core concepts from benchmarks and expanding them into diverse, knowledge-rich texts. By focusing on capability-level performance rather than benchmark-specific metrics, SuperValid’s loss correlates strongly and stably with downstream benchmark results across a wide range of models, scales, and training data distributions. This training‑free metric can be computed during training, enabling model selection, early stopping, and scaling decisions without the need for benchmark evaluation.

By Quanen Sun, Changxin Tian, Ke Shi, Cai Chen, Cunyin Peng, Jia Liu, Kunlong Chen, Zhiqiang Zhang, Jun Zhou
arXiv AI
Sep 16

Autonomous Assessment of Generalizability of AI Agent Capabilities

The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.

By Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava
arXiv AI
Aug 25

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

The paper introduces the Capability-Driven Multimodal Scaling Law, a cross-family framework that predicts vision-language model (VLM) benchmark accuracy from a low-dimensional textual capability score extracted via PCA. By training over 150 VLMs on 34 large language models across seven families, the authors demonstrate that the law accurately extrapolates transfer rates from 8B to 72B‑parameter backbones, predicts full training trajectories, and generalizes to unseen model families. The study also reveals actionable insights, such as certain textual benchmarks negatively correlating with multimodal performance and base LLMs outperforming instruction-tuned counterparts as VLM backbones due to higher absorption rates.

By Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai