The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.
By Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava
The paper introduces a formal framework for measuring AI propensities—tendencies of models to exhibit particular behaviours—using a bilogistic formulation that identifies an "ideal band" of success probability. It estimates the limits of this band with task‑agnostic rubrics and applies the method to six families of LLMs, showing how shifts in propensity affect task performance. The study finds that propensity estimates from one benchmark predict behaviour on held‑out tasks and that combining propensity with capability metrics yields stronger predictive power than either alone.
By Daniel Romero-Alvarado, Fernando Mart\'inez-Plumed, Lorenzo Pacchiardi, Hugo Save, Siddhesh Milind Pawar, Behzad Mehrbakhsh, Pablo Antonio Moreno Casares, Ben Slater, Paolo Bova, Peter Romero, Zachary R. Tidler, Jonathan Prunty, Luning Sun, Jose Hernandez-Orallo
arXiv:2607. 27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents.
By Peter Tisnikar, Maja Swieczkowska, Benteng Ma, Gerard Canal, Matteo Leonetti
arXiv:2608. 08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge.
By Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu
arXiv:2606. 06924v1 Announce Type: new Abstract: Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers.
By Guannan Lai, Haoran Hu, Long Chen, Zhenguo Li, Han-Jia Ye
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit.