Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

13,478 stories · RSS feed

arXiv Machine Learning
Aug 14

History-informed Lagrangian Neural Networks

arXiv:2608. 13215v1 Announce Type: new Abstract: Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and trajectory-specific physical properties must be inferred simultaneously.

By Tianshuo Zhang, Xianglei Xing, Wenzhe Zhai, Jia Gao, He Cao
arXiv Machine Learning
Aug 14

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

arXiv:2608. 12773v1 Announce Type: cross Abstract: Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day.

By Ebenezer Tarubinga
arXiv Machine Learning
Aug 14

SeBA: Semi-supervised few-shot learning via Separated-at-Birth Alignment for tabular data

arXiv:2605. 08519v2 Announce Type: replace Abstract: Learning from scarce labeled data with a larger pool of unlabeled samples, known as semi-supervised few-shot learning (SS-FSL), remains critical for applications involving tabular data in domains like medicine, finance, and science.

By Kacper Jurek, Wojciech Batko, Marek \'Smieja, Marcin Przewi\k{e}\'zlikowski
arXiv Machine Learning
Aug 14

TabH2O: A Unified Foundation Model for Tabular Prediction

arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.

By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati
arXiv AI
Aug 14

DiG-bench: Discovery in Games

arXiv:2608. 12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.

By Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, J\"urgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths, Zeb Kurth-Nelson, James C. R. Whittington
arXiv AI
Aug 14

Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

arXiv:2608. 13108v1 Announce Type: new Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decision contexts.

By Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu