Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,477 stories · RSS feed

arXiv AI
Jun 30

Diversity is the Strength of the AI Crowd

arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.

By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
arXiv AI
Jun 30

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.

By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
arXiv AI
Jun 30

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

arXiv:2606. 28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memory methods either truncate aggressively and lose non-local evidence or retain boilerplate that degrades decision quality.

By Mellow Baixuan Chen, Xiangguo Sun
arXiv AI
Jun 30

An AI agent for treatment reasoning over a biomedical tool universe

arXiv:2606. 28692v1 Announce Type: new Abstract: Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy.

By Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik
arXiv AI
Jun 30

The Human Creativity Benchmark

arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.

By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh