Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,030 stories · RSS feed

arXiv Machine Learning
Aug 11

Hyperbolic Multimodal Continual Learning

arXiv:2608. 09572v1 Announce Type: new Abstract: Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities.

By Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King
arXiv Machine Learning
Aug 11

Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy

arXiv:2608. 07580v1 Announce Type: cross Abstract: In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible Infrared Imaging Spectrometer (AVIRIS-3).

By William R. Keely, Philip G. Brodrick, Katherine Mistick, Adam Chlus, Robert O. Green, Philip E. Dennison
arXiv AI
Aug 11

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

arXiv:2608. 09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations.

By Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
arXiv Machine Learning
Aug 11

Analogical Learning for Cross-Scenario Generalization: Framework and Application to Intelligent Localization

arXiv:2504. 08811v3 Announce Type: replace Abstract: Modern learning systems often struggle with joint learning across diverse scenarios and immediate adaptation to new ones, because they rely heavily on the scenario-dependent absolute data-label representations.

By Zirui Chen, Hongning Ruan, Zhaoyang Zhang, Ziqing Xing, Ridong Li, Zhaohui Yang, M\'erouane Debbah
arXiv Machine Learning
Aug 11

SPD Learn: A Geometric Deep Learning Python Library for Neural Decoding Through Trivialization

arXiv:2602. 22895v2 Announce Type: replace-cross Abstract: Implementations of symmetric positive definite (SPD) matrix-based neural networks for neural decoding remain fragmented across research codebases and Python packages.

By Bruno Aristimunha, Ce Ju, Antoine Collas, Florent Bouchard, Ammar Mian, Bertrand Thirion, Sylvain Chevallier, Reinmar Kobler