arXiv AI
3d ago

Disclosure-Gated User Simulation for Companion-Agent Evaluation

The paper introduces a disclosure‑gated user simulator for companion‑agent evaluation, where information release is conditioned on the agent’s behavior through a five‑level gate system. The authors train the simulator on synthetic and real data, audit its performance, and demonstrate that it preserves ranking order and score stability across 12 tested systems. Their released simulator achieves a 0.993 correlation with the original benchmark’s simulator and outperforms prompting a frontier model, which only shifts scores upward without affecting rankings.

By Yao Liu, Yu He
arXiv Machine Learning
Jun 25

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

arXiv:2606. 25760v1 Announce Type: new Abstract: Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions.

By Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
arXiv AI
Jun 3

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.

By Alexander Apartsin, Yehudit Aperstein