Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,528 stories · RSS feed

arXiv Machine Learning
Aug 4

NORi: An ML-Augmented Ocean Boundary Layer Parameterization

arXiv:2512. 04452v3 Announce Type: replace-cross Abstract: NORi is a machine learning (ML) parameterization of ocean boundary layer turbulence that is physics-based and augmented with neural networks.

By Xin Kai Lee, Ali Ramadhan, Andre Souza, Gregory LeClaire Wagner, Simone Silvestri, John Marshall, Raffaele Ferrari
arXiv Machine Learning
Aug 4

Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language

arXiv:2605. 15607v2 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve high pass rates on code generation benchmarks, yet whether they can transfer this ability to languages absent from pretraining remains poorly understood.

By Vinayshekhar Bannihatti Kumar, Disha Makhija, Manoj Ghuhan Arivazhagan, Rashmi Gangadharaiah
arXiv Machine Learning
Aug 4

Tevatron Meets Megatron: Expert-Parallel LLM Reranker Training on an Academic Budget

arXiv:2608. 00916v1 Announce Type: cross Abstract: Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups.

By Zhichao Xu, Xueguang Ma, Shengyao Zhuang, Luyu Gao, Wenqian Ye, Yu Wang, Jamie Callan, Jimmy Lin
arXiv Machine Learning
Aug 4

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

arXiv:2608. 01202v1 Announce Type: cross Abstract: Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-ranging advantages in agriculture field for both pre-harvest and post-harvest management.

By Ahmed Baha Ben Jmaa, Faten Chaieb, Anna Fabija\'nska
arXiv Machine Learning
Aug 4

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

arXiv:2608. 01821v1 Announce Type: cross Abstract: Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual conditioning a substantial recurring inference cost.

By Yongkang Zhou, Xiang Xia, Cheng Yan, Fan Xu, Wuyang Zhang