Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

19,504 stories · RSS feed

arXiv Machine Learning
Jun 26

CALIBURN: Operationally Calibrated Streaming Intrusion Detection with Regime-Dependent Conformal Risk Control

arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.

By Michel A. Youssef
arXiv Machine Learning
Jun 26

Tractography-Driven Synthetic Data Generation for Fiber Bundle Segmentation in Tracer Histology

arXiv:2606. 26898v1 Announce Type: cross Abstract: Diffusion MRI (dMRI) tractography enables non-invasive reconstruction of white-matter pathways, but its accuracy is fundamentally limited by indirect, low-resolution measurements of axonal organization.

By Kyriaki-Margarita Bintsi, Sparsh Makharia, Ya\"el Balbastre, Joselyn Romero Avila, Julia F. Lehman, Suzanne N. Haber, Anastasia Yendiki
arXiv AI
Jun 26

Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing

arXiv:2603. 30014v2 Announce Type: replace-cross Abstract: The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LHC), has evolved into a robust platform for orchestrating large-scale workflows across distributed computing resources.

By Derek Anderson, Amit Bashyal, Markus Diefenthaler, Cristiano Fanelli, Wen Guan, Tanja Horn, Alex Jentsch Meifeng Lin, Tadashi Maeno, Kei Nagai, Hemalata Nayak, Connor Pecar, Karthik Suresh, Fang-Ying Tsai, Anselm Vossen, Tianle Wang, Torre Wenaus
arXiv AI
Jun 26

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

arXiv:2606. 26458v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG).

By Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma