Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

13,227 stories · RSS feed

arXiv AI
Aug 13

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

arXiv:2608. 12307v1 Announce Type: cross Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods.

By Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
arXiv AI
Aug 13

Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection

arXiv:2512. 13374v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether their internal representations capture problem structure or algorithmic behavior.

By Francesca Da Ros, Luca Di Gaspero, Kevin Roitero
arXiv AI
Aug 13

NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

arXiv:2608. 12197v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning.

By Jiarui Ma, Jianghan Wang, Yuheng Ma, Ziyi Zhuang, Xiaoguang Liu
arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv Machine Learning
Aug 13

Generative Learning for Quantum Measurement Design

arXiv:2608. 11396v1 Announce Type: cross Abstract: Extracting quantum information from a quantum state is a fundamental task of quantum computation, often requiring the estimation of many non-commuting observables under a finite measurement budget.

By Jun Dai, Olivier Nahman-L\'{e}vesque, Guillaume Rabusseau, Hong-Ye Hu, Cunlu Zhou
arXiv AI
Aug 13

Teaching agentic AI to learn expert reasoning for rare disease diagnosis

arXiv:2606. 16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.

By Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
arXiv AI
Aug 13

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

arXiv:2608. 11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression.

By Angelo Nardone, Paolo Ferragina