Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,907 stories · RSS feed

arXiv Machine Learning
Aug 4

How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

arXiv:2608. 01454v1 Announce Type: cross Abstract: Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols.

By Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen
arXiv Machine Learning
Aug 4

Physics constraints and response validation in discrete-time reduced-order modeling: from idealized turbulent systems to climate dynamics

arXiv:2602. 13847v5 Announce Type: replace-cross Abstract: A central challenge across science and engineering is to build data-driven reduced-order models of turbulent dynamical systems that reproduce stationary statistics, predict responses to external perturbations, and remain practical for real-world applications.

By Fabrizio Falasca, Laure Zanna
arXiv Machine Learning
Aug 4

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

arXiv:2603. 11804v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Delyan Boychev (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")