arXiv AI

Capabilities Ain't All You Need: Measuring Propensities in AI

The paper introduces a formal framework for measuring AI propensities—tendencies of models to exhibit particular behaviours—using a bilogistic formulation that identifies an "ideal band" of success probability. It estimates the limits of this band with task‑agnostic rubrics and applies the method to six families of LLMs, showing how shifts in propensity affect task performance. The study finds that propensity estimates from one benchmark predict behaviour on held‑out tasks and that combining propensity with capability metrics yields stronger predictive power than either alone.

arXiv AI
Aug 6

Item Response Theory for AI Safety

arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.

By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
arXiv AI
Sep 16

Autonomous Assessment of Generalizability of AI Agent Capabilities

The paper introduces Monte Carlo Query Search (MCQS), an active query‑synthesis method for learning symbolic stochastic capability models of black‑box AI agents. MCQS treats capability evaluation as an active learning problem over policies, using Monte Carlo tree search to generate queries that distinguish between pessimistic and optimistic capability hypotheses. Experiments demonstrate that MCQS learns accurate capability models more efficiently than baseline strategies, enabling systematic characterization of agent capability boundaries with fewer interactions.

By Daniel Bramblett, Rushang Karia, Adrian Ciotinga, Pulkit Verma, YooJung Choi, Siddharth Srivastava
arXiv Machine Learning
Sep 22

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

The paper introduces SCOPE, a method that post‑trains computer‑use agents to balance task completion with safety by conditioning actions on environmental risk. It combines supervised fine‑tuning on three trajectory types—capability demonstrations, safe continuations, and explicit refusals—followed by reinforcement learning to improve performance. Experiments starting from Qwen3.5‑9B show that SCOPE‑RL achieves high task success and attack‑avoidance rates, outperforming other agents on OSWorld and OS‑BLIND benchmarks.

By Zeyu Kang, Zhenyun Yin, Yang Zhang, Shan He, Shanzhe Lei, Yanjiu Zhong, Xinquan Chen, Yuhong Wang
arXiv AI
Jun 10

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.

By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
arXiv AI
6d ago

Efficient Safety Benchmarking via Item Response Theory

The paper demonstrates that Item Response Theory (IRT) can uncover meaningful structure in safety benchmarks for language models, allowing adaptive item selection to approximate full benchmark rankings with Spearman’s ρ > 0.90 while cutting evaluation costs by at least 80% and up to 99.9% on some suites. It also proposes a static method to extract a small, informative subset of items that can be reused across models, achieving 80–99.8% cost savings. These findings show that psychometric techniques can make safety evaluation more efficient without sacrificing ranking accuracy.

By Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri, Diogo Cruz