Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,270 stories · RSS feed

arXiv AI
Jul 7

Phase-Preserving Trimodal Transformer for Tropical Forest Biomass Estimation Using Optical and PolInSAR Data

arXiv:2607. 03663v1 Announce Type: cross Abstract: The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily due to the saturation of Synthetic Aperture Radar (SAR) signals in high-density areas and persistent cloud cover affecting optical imagery.

By Luiz Felipe Parente Santiago (Institute of Computing, Brazilian Army Research Institute in the Amazon), Rosiane Rodrigues de Freitas (Institute of Computing), Daniel Rodrigues dos Santos (Military Institute of Engineering), Felipe Ferrari (Military Institute of Engineering)
arXiv Machine Learning
Jul 7

GeoFlow: Geo-Aware Modeling of Inter-Area Relationships in Origin-Destination Flow Prediction and Generation

arXiv:2607. 05257v1 Announce Type: new Abstract: Origin-destination (OD) flow modeling underpins urban planning and mobility analysis, but prevailing graph-based methods often neglect salient geographic attributes, limiting their ability to model long-range and multi-area dependencies.

By Zherui Huang, Guanjie Zheng, Hao Xue, Linghe Kong
arXiv AI
Jul 7

Deriving Neural Scaling Laws from the statistics of natural language

arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.

By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
arXiv Machine Learning
Jul 7

SMART: A Machine Learning and Monte Carlo Framework for Rapid Analysis of Stochastic Transistor Aging and Process Variation in Digital Circuits

arXiv:2607. 05187v1 Announce Type: new Abstract: As CMOS technology scales into the deep nanometer regime, digital circuit reliability is increasingly threatened by the combined stochastic effects of Bias Temperature Instability (BTI) and Process Variation (PV).

By Arash Esshaghi, Siavash Es'haghi, Gholamreza Shahabadi, Alireza Moradi
arXiv AI
Jul 7

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

arXiv:2607. 03126v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, Yanhua Cheng, Peng Jiang, Yadong Mu