Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,477 stories · RSS feed

arXiv AI
Jun 30

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

arXiv:2606. 30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time.

By Rahul Khedar, Mayank Malhotra, Avinash Karn, Mouli V, Prakhar Mehrotra
arXiv AI
Jun 30

The Human Creativity Benchmark

arXiv:2606. 30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved.

By Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh
arXiv AI
Jun 30

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.

By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
arXiv Machine Learning
Jun 30

fev-bench: A Realistic Benchmark for Time Series Forecasting

arXiv:2509. 26468v3 Announce Type: replace Abstract: Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models.

By Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen, Lorenzo Stella, Nick Erickson, Pablo Guerron, Michael Bohlke-Schneider, Yuyang Wang
arXiv Machine Learning
Jun 30

Atompack: A Storage and Distribution Layer for Read-Heavy Atomistic ML Training Datasets

arXiv:2606. 29975v1 Announce Type: new Abstract: Atomistic machine learning datasets are increasingly used for training: large immutable snapshots are read repeatedly, shuffled across epochs, staged across clusters' storage systems, and republished as reusable scientific artifacts.

By Ali Ramlaoui, Daniel T. Speckhard, Sagar Pal, Fragkiskos D. Malliaros, Alexandre Duval, Victor Schmidt
arXiv Machine Learning
Jun 30

MALOQ: Massively Accelerated Learning of Operators for Quantum Transport

arXiv:2606. 28911v1 Announce Type: new Abstract: Machine-learned (ML) operator models can be trained to predict density functional theory (DFT) Hamiltonian/density matrices at significantly reduced computational cost, thus extending electronic-structure calculations to previously unfeasible scales.

By Manasa Kaniselvan, Alexander Maeder, Denghui Lu, Alexandros Nikolaos Ziogas, Mathieu Luisier