Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

17,521 stories · RSS feed

arXiv Machine Learning
Jul 9

Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector

arXiv:2607. 07382v1 Announce Type: new Abstract: Fast Radio Bursts (FRBs) are millisecond-duration radio transients whose automated detection increasingly relies on highly specialized deep learning models.

By Raiff H. Santos, Amilcar R. Queiroz, Tharcisyo S. S. Duarte, K. E. L. de Farias, Rafael A. Batista
arXiv Machine Learning
Jul 9

Geometric--Nongeometric Optimizer Calculus: A Modular Language for Reachable Gradient Methods

arXiv:2607. 07206v1 Announce Type: new Abstract: Adaptive optimizers mix several mechanisms: a metric or preconditioner maps gradients to descent directions, while estimation, memory, step-size control, constraints, stochasticity, target modification, and discretization determine which directions are available and how they are used.

By Zavier Li
arXiv AI
Jul 9

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

arXiv:2507. 05116v5 Announce Type: replace-cross Abstract: Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language.

By Juyi Lin, Amir Taherin, Arash Akbari, Arman Akbari, Lei Lu, Guangyu Chen, Taskin Padir, Xiaomeng Yang, Weiwei Chen, Yiqian Li, Xue Lin, David Kaeli, Pu Zhao, Yanzhi Wang