Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,048 stories · RSS feed

arXiv AI
Aug 5

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

arXiv:2608. 03597v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality.

By Amirhossein Taleshinosrati, Yangyang Wang, Atitaya Phoemsuk, Vahid Abolghasemi, Naser Hossein Motlagh, Sadasivan Puthusserypady, Daniel Teichmann, Abdolrahman Peimankar
arXiv Machine Learning
Aug 5

Maglev: Sliding Recurrent Memory

arXiv:2608. 02870v1 Announce Type: new Abstract: We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training.

By Bo Liu, Qiang Liu
arXiv AI
Aug 5

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

arXiv:2510. 05159v5 Announce Type: replace-cross Abstract: While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain.

By L\'eo Boisvert, Abhay Puri, Chandra Kiran Reddy Evuru, Nazanin Sepahvand, Nicolas Chapados, Quentin Cappart, Jason Stanley, Alexandre Lacoste, Krishnamurthy Dj Dvijotham, Alexandre Drouin
arXiv Machine Learning
Aug 5

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

arXiv:2603. 25857v3 Announce Type: replace Abstract: The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction.

By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Christian Feiler, Roland C. Aydin
arXiv AI
Aug 5

Rubrics as Privileged Information for Open-Ended Generation

arXiv:2608. 02948v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations.

By Deepika Bablani, Ajay Gupta, Wanming Chen
arXiv Machine Learning
Aug 5

Trajectory inference via Acceleration Matching

arXiv:2608. 03916v1 Announce Type: new Abstract: Trajectory inference is a fundamental problem in many scientific domains: given a collection of unpaired snapshots of observations at discrete time points, the goal is to generate smooth trajectories that best resemble and interpolate the data.

By Bartolo Dazzini, Giovanni Conforti, Alain Durmus, Aram-Alexandre Pooladian