Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,220 stories · RSS feed

arXiv Machine Learning
Aug 4

Physics constraints and response validation in discrete-time reduced-order modeling: from idealized turbulent systems to climate dynamics

arXiv:2602. 13847v5 Announce Type: replace-cross Abstract: A central challenge across science and engineering is to build data-driven reduced-order models of turbulent dynamical systems that reproduce stationary statistics, predict responses to external perturbations, and remain practical for real-world applications.

By Fabrizio Falasca, Laure Zanna
arXiv Machine Learning
Aug 4

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

arXiv:2603. 11804v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Delyan Boychev (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
arXiv Machine Learning
Aug 4

Semi-MedRef: Semi-Supervised Medical Referring Image Segmentation with Cross-Modal Alignment

arXiv:2605. 15720v2 Announce Type: replace-cross Abstract: Medical referring image segmentation (MRIS) predicts lesion masks from medical images and natural-language referring expressions, but acquiring paired pixel-level annotations and referring texts is costly.

By Yuchen Li, Ziru Wei, Zhen Zhao, Yi Liu, Luping Zhou