Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

22,912 stories · RSS feed

arXiv AI
Jun 10

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.

By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
arXiv AI
Jun 10

IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts

arXiv:2606. 09908v1 Announce Type: cross Abstract: Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evaluation.

By Ayana Hussain, Soumya Sharma, Golnoosh Farnadi, Nicholas Vincent, H\'eber Hwang Arcolezi, Ulrich A\"ivodji
arXiv AI
Jun 10

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

arXiv:2606. 10267v1 Announce Type: cross Abstract: Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals executed by low-level VLA controllers.

By Jiaheng Hu, Mohit Shridhar, Caden Lu, Dhruv Shah, Hao-Tien Lewis Chiang, Jie Tan, Annie Xie
arXiv Machine Learning
Jun 10

The hyper-scaled NLP bound for maximum-entropy remote sampling

arXiv:2601. 20970v3 Announce Type: replace-cross Abstract: The maximum-entropy remote sampling problem (MERSP) is to select a subset of $s$ random variables from a set of $n$ random variables, so as to maximize the information concerning a set of target random variables that are not directly observable.

By Gabriel Ponte, Marcia Fampa, Jon Lee
arXiv AI
Jun 10

GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support

arXiv:2604. 12306v3 Announce Type: replace-cross Abstract: Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-specific climate knowledge and grounded interaction with geospatial and forecasting tools.

By Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan, Muhammad Haris Khan
arXiv AI
Jun 10

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

arXiv:2606. 10084v1 Announce Type: cross Abstract: This work presents a divide-and-conquer modeling strategy for the CTF-4-Science Lorenz benchmark, which evaluates chaotic-system prediction across twelve hidden scores and five scenario families: clean forecasting, noisy reconstruction, noisy-input forecasting, few-shot learning, and parametric generalization.

By Shundong Li