Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,285 stories · RSS feed

arXiv Machine Learning
Jul 2

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video

arXiv:2603. 16432v3 Announce Type: replace-cross Abstract: Unsupervised physical parameter estimation from video lacks a common benchmark: existing methods evaluate on non-overlapping synthetic data, the sole real-world dataset is restricted to single-body systems, and no established protocol addresses governing-equation identification.

By Rasul Khanbayov, Mohamed Rayan Barhdadi, Erchin Serpedin, Hasan Kurban
arXiv AI
Jul 2

EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology

arXiv:2607. 01039v1 Announce Type: cross Abstract: Therapy-induced cardiotoxicity is the leading non-oncological cause of treatment interruption in breast cancer patients, yet early, automated risk stratification from routine cardiac imaging remains an unsolved problem.

By Grigorios Kalliatakis, Georgia Karanasiou, Georgios Manikis, Manolis Tsiknakis, Dimitrios Fotiadis, Dorothea Tsekoura, Kalliopi Keramida, Vasileios Bouratzis, Lampros Lakkas, Katerina Naka, Andri Papakonstantinou, Anastasia Constantinidou, Kostas Marias
arXiv AI
Jul 2

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

arXiv:2607. 01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task.

By Brett Reynolds
arXiv AI
Jul 2

LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data

arXiv:2607. 00733v1 Announce Type: cross Abstract: Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of underlying mechanisms, which is particularly valuable in clinical settings.

By Hanning Yang, Meropi Karakioulaki, Lennart Purucker, Tim Litwin, Cristina Has, Moritz Hess
arXiv Machine Learning
Jul 2

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

arXiv:2605. 04035v3 Announce Type: replace-cross Abstract: We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups.

By Evangelos Ntavelis, Sean Wu, Mohamad Shahbazi, Fabio Maninchedda, Dmitry Kostiaev, Artem Sevastopolsky, Vittorio Megaro, Trevor Phillips, Alejandro Blumentals, Shridhar Ravikumar, Mehak Gupta, Reinhard Knothe, Jeronimo Bayer, Matthias Vestner, Simon Schaefer, Thomas Etterlin, Christian Zimmermann, Alexey Artemov, Mathias Deschler, Peter Kaufmann, Stefan Brugger, Sebastian Martin, Brian Amberg, Tom Runia