Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,477 stories · RSS feed

arXiv Machine Learning
Jun 30

Momentum Guidance: Plug-and-Play Guidance for Flow Models

arXiv:2602. 20360v2 Announce Type: replace Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used in their vanilla conditional form: in image generation, samples without guidance often appear diffuse and lack fine-grained detail.

By Runlong Liao, Jian Yu, Baiyu Su, Chi Zhang, Lizhang Chen, Qiang Liu
arXiv AI
Jun 30

CaveAgent: Transforming LLMs into Stateful Runtime Operators

arXiv:2601. 01569v4 Announce Type: replace Abstract: LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that struggle with long-horizon tasks due to fragile multi-turn dependencies and context drift.

By Maohao Ran, Zhenglin Wan, Cooper Lin, Yanting Zhang, Hongyu Xin, Hongwei Fan, Yibo Xu, Beier Luo, Yaxin Zhou, Wangbo Zhao, Lijie Yang, Lang Feng, Fuchao Yang, Jingxuan Wu, Yiqiao Huang, Chendong Ma, Yusen Huang, Dailing Jiang, Jianbo Deng, Sirui Han, Yang You, Bo An, Yike Guo, Jun Song
arXiv AI
Jun 30

Reinforcement Learning for Software Vulnerability Analysis: A Systematic Review with Emphasis on C/C++ Source Code and Static Analysis

arXiv:2606. 28403v1 Announce Type: cross Abstract: Vulnerability detection in C/C++ software remains a major security challenge due to code complexity, manual memory management, and the limitations of traditional static analysis.

By Bruno Caro-V\'asquez, Carola Figueroa-Flores, Gast\'on Marquez
arXiv AI
Jun 30

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval

arXiv:2606. 28379v1 Announce Type: cross Abstract: We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents must be applied efficiently without breaking cross-references or semantic consistency.

By Mike Hang Wang, Utkarsh Garg, Reza Davari, Huitian Jiao, Hao Cheng, Baolin Peng, Tao Ge, Si-Qing Chen
arXiv AI
Jun 30

Evolutional Math: Cross-Validated Island-Model Genetic Programming for Interpretable Symbolic Regression on Small, Wide Datasets

arXiv:2606. 28381v1 Announce Type: cross Abstract: Symbolic regression via genetic programming routinely fails on small, wide datasets - a regime common in clinical-trial monitoring, biostatistics, and engineering pilot studies - by converging on bloated, overfit expressions that exploit correlation rather than prediction.

By Artem Andrianov (Cyntegrity Germany GmbH, Hofheim am Taunus, Germany)
arXiv Machine Learning
Jun 30

CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification

arXiv:2605. 16401v2 Announce Type: replace-cross Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environmental impact, and a "one-size-fits-all" approach that ignores varying sample complexity.

By Mikael Turkoglu, Tim Bary, Vincent Thielens, Manon Dausort, Beno\^it Macq
arXiv Machine Learning
Jun 30

Audio-Visual Continual Test-Time Adaptation without Forgetting

arXiv:2602. 18528v2 Announce Type: replace Abstract: Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, which hampers online cross-modal learning and eventually leads to poor accuracy.

By Sarthak Kumar Maharana, Akshay Mehra, Bhavya Ramakrishna, Yunhui Guo, Guan-Ming Su