Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,477 stories · RSS feed

arXiv AI
Jun 30

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

arXiv:2406. 08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate causal relationships that go beyond pairwise correlations largely unmeasured.

By Zineb Senane, Axel Karlsson, Lele Cao, Oleg Smirnov, Cheng Zhang, Sahar Asadi, Hedvig Kjellstr\"om, Gustav Eje Henter, Ruibo Tu
arXiv Machine Learning
Jun 30

CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification

arXiv:2605. 16401v2 Announce Type: replace-cross Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environmental impact, and a "one-size-fits-all" approach that ignores varying sample complexity.

By Mikael Turkoglu, Tim Bary, Vincent Thielens, Manon Dausort, Beno\^it Macq
arXiv AI
Jun 30

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

arXiv:2606. 29502v1 Announce Type: new Abstract: Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another.

By Songjun Tu, Chengdong Xu, Qichao Zhang, Yiwen Ma, Yaocheng Zhang, Linjing Li, Dong Li, Xiangyuan Lan, Dongbin Zhao