Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

22,530 stories · RSS feed

arXiv AI
Jun 9

FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation

arXiv:2512. 15116v2 Announce Type: replace-cross Abstract: Multivariate time series imputation is fundamental in applications such as healthcare, traffic forecasting, and biological modeling, where sensor failures and irregular sampling lead to pervasive missing values.

By Runze Li, Hanchen Wang, Wenjie Zhang, Binghao Li, Yu Zhang, Xuemin Lin, Ying Zhang
arXiv AI
Jun 9

Training-Inference Kernel Contracts: Bounding Divergence in Post-Training and Deployment

arXiv:2606. 07581v1 Announce Type: cross Abstract: A modern post-training pipeline often writes one symbol for its policy, pi_theta, while evaluating it through two different programs: a training kernel optimized for autograd and an inference kernel optimized for low-precision, fused, dynamically batched serving.

By Bruce Changlong Xu, Lan Wu
arXiv AI
Jun 9

From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing

arXiv:2606. 08932v1 Announce Type: cross Abstract: Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but silently drops nested exceptions or counter-exceptions, producing outputs that appear compliant yet break on important edge cases.

By Jian Chen, Siyuan Li, Chucheng Wan, Zixuan Yuan