Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,014 stories · RSS feed

arXiv Machine Learning
Aug 11

Full-Feature versus Limited-Input Machine Learning for Residential Energy Estimation: A Comparative Analysis of RECS and ResStock Under Realistic Input Constraints

arXiv:2608. 09255v1 Announce Type: new Abstract: Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data are available.

By Aditya Ramnarayan, Fatih Evren, Patti Gunderson, Samuel Rosenberg
arXiv Machine Learning
Aug 11

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

arXiv:2608. 07895v1 Announce Type: cross Abstract: Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction.

By Simon Holk, Ryosuke Takanami, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Yueh-Hua Wu, Kei Ota
arXiv Machine Learning
Aug 11

Failure-Mechanism Transferability of Cumulative-Damage Features for Health State Estimation of SiC Power Modules

arXiv:2608. 08365v1 Announce Type: cross Abstract: Data-driven health-state estimators for SiC (Silica-Carbide) power modules typically report their performance on a single accelerated-aging campaign, and how that performance transfers to a different failure mechanism is rarely tested.

By Mattia Scarpa, Evgeny Kusmenko, Francesco Toso, Mattia Bruschetta, Ruggero Carli, Simon Achatz
arXiv Machine Learning
Aug 11

Eikonal Regularisation in Physics-Informed Neural Networks for Three-Dimensional Level-Set Advection: Transferability of Two-Dimensional Design Principles

arXiv:2608. 08322v1 Announce Type: cross Abstract: Physics-informed neural networks applied to the level-set formulation of interface advection commonly augment the residual and initial-condition losses with an eikonal regulariser, penalising the deviation of $\|\nabla\phi\|$ from unity.

By Muhammad Akbar Khan
arXiv Machine Learning
Aug 11

Kernel Methods for Refined Prophet Inequalities

arXiv:2608. 08662v1 Announce Type: cross Abstract: The single-selection prophet inequality is a canonical Bayesian online selection problem in which independent nonnegative values arrive sequentially and the decision-maker must irrevocably select at most one.

By Patrick Loiseau, Mathieu Molina, Vianney Perchet, Sebastian Perez-Salazar, Victor Verdugo
arXiv Machine Learning
Aug 11

Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations

arXiv:2608. 09053v1 Announce Type: cross Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence.

By Hongxiang Gao, He-yang Xu, Yuwen Li, Minghui Zhao, Zhipeng Cai, Xingyao Wang, Chenxi Yang, Jianqing Li, Chengyu Liu
arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv Machine Learning
Aug 11

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

arXiv:2608. 09209v1 Announce Type: cross Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs.

By Chidaksh Ravuru, Shashank Srivastava
arXiv Machine Learning
Aug 11

Learning to Triage Vulnerability Reports from Program Analysis: An Empirical Study in Node.js

arXiv:2510. 20739v2 Announce Type: replace-cross Abstract: Program analysis tools often produce large volumes of candidate vulnerability reports that require costly manual review, creating a practical challenge: how can security analysts prioritize the reports most likely to be true vulnerabilities?

By Ronghao Ni, Aidan Z. H. Yang, Min-Chien Hsu, Nuno Sabino, Limin Jia, Ruben Martins, Darion Cassel, Kevin Cheang