Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,829 stories · RSS feed

arXiv Machine Learning
Jul 27

Unbiased Open World Regularization for Fair Self-Supervised Learning

arXiv:2607. 22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset.

By L{\'e}o Nicollier (CB, ATT), Marc Pic (ATT), Pablo Mus{\'e} (CB, IFUMI), Enric Meinhardt-Llopis (CB), Gabriele Facciolo (CB)
arXiv Machine Learning
Jul 27

Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes

arXiv:2607. 22508v1 Announce Type: new Abstract: Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks.

By Athanasios Papastathopoulos-Katsaros, Steven T. Lee, Lin Yao, Ajay Thomas, Junseok Park, Matthew J. McGinley, Zhandong Liu
arXiv Machine Learning
Jul 27

Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification

arXiv:2607. 21186v1 Announce Type: cross Abstract: Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning.

By Guillermo Rubi\~nos Rodr\'iguez, Mart\'in Ottavianelli, Mateo Alonso, Gonzalo Bl\'azquez Gil, Boris-Stephan Rauchmann, Pablo D\'iez-Valle, Sergio Altares-L\'opez
arXiv Machine Learning
Jul 27

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

arXiv:2607. 21610v1 Announce Type: cross Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available.

By Miaobo Hu, Xiaobo Guo, Shuhao Hu, Bokun Wang, Rui Chen, Xin Wang, Daren Zha, Jun Xiao
arXiv Machine Learning
Jul 27

CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography

arXiv:2607. 22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols.

By Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo\'nski, Tomasz Figatowski, Natalia Zieli\'nska