Monte Carlo Query Search: Active Capability Assessment of AI Agents
arXiv:2512. 16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making.
The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.
arXiv:2512. 16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making.
The paper introduces a formal framework for measuring AI propensities—tendencies of models to exhibit particular behaviours—using a bilogistic formulation that identifies an "ideal band" of success probability. It estimates the limits of this band with task‑agnostic rubrics and applies the method to six families of LLMs, showing how shifts in propensity affect task performance. The study finds that propensity estimates from one benchmark predict behaviour on held‑out tasks and that combining propensity with capability metrics yields stronger predictive power than either alone.
The paper introduces a capability manifold, a multidimensional framework that maps downstream capabilities—such as reasoning, retrieval, planning, and adaptation—to pre‑training, post‑training, and test‑time resources via bounded scaling functions. It provides analytical Jacobians to quantify how sensitive each capability is to changes in resources and their interactions. By embedding existing Kaplan‑ and Chinchilla‑type scaling laws and test‑time compute into this manifold, the authors demonstrate that these scaling relationships can be unified as trajectories on a common capability manifold.
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insuffi...
arXiv:2609.27869v1 Announce Type: new Abstract: Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution. Existing designs typ...
The paper introduces SCOPE, a method that post‑trains computer‑use agents to balance task completion with safety by conditioning actions on environmental risk. It combines supervised fine‑tuning on three trajectory types—capability demonstrations, safe continuations, and explicit refusals—followed by reinforcement learning to improve performance. Experiments starting from Qwen3.5‑9B show that SCOPE‑RL achieves high task success and attack‑avoidance rates, outperforming other agents on OSWorld and OS‑BLIND benchmarks.
arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.
arXiv:2606. 05296v1 Announce Type: new Abstract: LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be controlled purely at test time.
arXiv:2604. 05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the target environment.
arXiv:2608. 00316v1 Announce Type: new Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic statistical priors.
arXiv:2602. 21556v2 Announce Type: replace Abstract: When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce a synthesized output.
arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...