Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

22,912 stories · RSS feed

arXiv Machine Learning
Jun 9

Hardware-aware Low-latency Quantum Compilation with Data-driven Lightweight Error Detection for Early Fault-Tolerant Systems

arXiv:2606. 07666v1 Announce Type: cross Abstract: Noisy intermediate-scale quantum (NISQ) processors are entering an early fault-tolerance regime where full quantum error correction carries prohibitive resource costs, yet lightweight error detection can meaningfully improve algorithmic success rates.

By Sumit Chongder (Indian Institute of Technology Jodhpur)
arXiv AI
Jun 9

Bayesian Selective Latent Inference for Wastewater-First Influenza Monitoring

arXiv:2606. 09433v1 Announce Type: new Abstract: Wastewater influenza surveillance can reveal community circulation before clinical reporting, but wastewater alone is not a fully identifiable proxy for human burden.

By Yixuan Zhang (Section of Health Data Science and AI, Department of Public Health, University of Copenhagen, Copenhagen, Denmark), Yang Song (Section of Health Data Science and AI, Department of Public Health, University of Copenhagen, Copenhagen, Denmark), Hao Wang (Rutgers University, New Brunswick, NJ, USA), Samir Bhatt (Section of Health Data Science and AI, Department of Public Health, University of Copenhagen, Copenhagen, Denmark, MRC Centre for Global Infectious Disease Analysis, Department of Infectious Disease Epidemiology, School of Public Health, Faculty of Medicine, Imperial College London, London, United Kingdom), Hengguan Huang (Section of Health Data Science and AI, Department of Public Health, University of Copenhagen, Copenhagen, Denmark, MRC Centre for Global Infectious Disease Analysis, Department of Infectious Disease Epidemiology, School of Public Health, Faculty of Medicine, Imperial College London, London, United Kingdom)
arXiv AI
Jun 9

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.

By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su
arXiv AI
Jun 9

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.

By Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru, Zhejun Cui, Weitong Lian, Qi Kang, Hangshuo Cao, Yichen Zhu, Yechi Liu, Kaixuan Wang, Yu-Jie Yuan, Chunwei Wang, Yu Zhang, Bo Dai
arXiv AI
Jun 9

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

arXiv:2606. 08376v1 Announce Type: cross Abstract: As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have grown in frequency and diversity.

By Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan