Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,285 stories · RSS feed

arXiv AI
Jul 1

Qualified Educational Capacity Planning under Heterogeneous Student Support Needs: A Synthetic Benchmark and Decision-Support Framework

arXiv:2606. 30650v1 Announce Type: cross Abstract: Educational support services often face a qualified-capacity problem: staff time is scarce, qualifications decay, new support needs can appear before anyone is prepared for them, and training consumes the same hours needed by current students.

By Carlos Eduardo Sanoja, Oscar Enrique Moreno Mayz
arXiv AI
Jul 1

Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

arXiv:2606. 31800v1 Announce Type: new Abstract: Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static supervision, where fixed prompts, rules, or reward models provide non-adaptive guidance throughout training.

By Xianda Zheng, Huan Gao, Meng-Fen Chiang, Michael Witbrock, Kaiqi Zhao, Shangyang Li
arXiv AI
Jul 1

The HydroGym Reinforcement Learning Platform for Fluid Dynamics

arXiv:2512. 17534v2 Announce Type: replace-cross Abstract: Modeling and controlling fluids is critical across science and engineering.

By Christian Lagemann, Sajeda Mokbel, Miro Gondrum, Mario R\"uttgers, Yuning Wang, Pol Su\'arez, Ludger Paehler, Deniz A. Bezgin, Aaron B. Buhendwa, Jared L. Callaham, Samuel Ahnert, Nicholas Zolman, Xiao Shao, Jean-Christophe Loiseau, Nikolaus Adams, Matthias Meinke, Wolfgang Schr\"oder, Kai Lagemann, Esther Lagemann, Ricardo Vinuesa, Steven L. Brunton
arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee