Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

16,065 stories · RSS feed

arXiv Machine Learning
Jul 22

Neural Kolmogorov Equations: Parallelizable Learning of Stochastic Dynamics under General Noise

arXiv:2607. 19173v1 Announce Type: new Abstract: Neural stochastic differential equations (SDEs) have emerged as powerful tools for learning noisy or stochastic dynamics directly from data; however, existing approaches largely assume uncoupled and continuous noise, limiting their applicability to realistic stochastic drivers, and often scale poorly in time, requiring expensive autoregressive training.

By Arthur Bizzi, Olga Fink
arXiv Machine Learning
Jul 22

Recti-Q: Feature-Space Rectification for Out-of-Distribution-Robust Quantized Perception in Edge Robotics

arXiv:2607. 18540v1 Announce Type: cross Abstract: Robotic perception pipelines increasingly rely on large vision backbones deployed on SWaP-constrained edge platforms, making post-training quantization (PTQ) attractive for real-time inference.

By Hamidreza Yaghoubi Araghi, Parastoo Pilevar, Ming C. Lin
arXiv AI
Jul 22

Robust Reasoning Benchmark

arXiv:2604. 08571v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting.

By Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko, Mark C. Jeffrey
arXiv AI
Jul 22

Now We Know? A Systematic Comparison of TerraMind and THOR

arXiv:2607. 18504v1 Announce Type: cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact?

By Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
arXiv Machine Learning
Jul 22

Enhanced NQS via Annealed Gradient Descent

arXiv:2607. 18865v1 Announce Type: cross Abstract: Neural quantum states offer expressive representations of quantum many-body wave functions, yet their practical accuracy can be limited by stochastic optimization rather than representational capacity.

By Shiwei Zhou, Yiming Huang, Xiao Yuan, Xiaoxia Cai
arXiv Machine Learning
Jul 22

Quantum Reservoir Computing: Recent Advances and Future Directions

arXiv:2607. 18552v1 Announce Type: cross Abstract: Quantum reservoir computing (QRC) uses the dynamics of a fixed or weakly tuned quantum system to transform temporal and sequential inputs into measured features, while training is typically confined to a classical readout.

By Shehbaz Tariq, Muhammad Talha, Arshid Ali, Muhammad Diyan, Symeon Chatzinotas