Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

12,559 stories · RSS feed

arXiv Machine Learning
4d ago

WANDR: A Benchmark for Wide and Deep Research

arXiv:2608. 14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents.

By Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma
arXiv Machine Learning
4d ago

Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference

arXiv:2608. 15282v1 Announce Type: new Abstract: Earth observation foundation models (EOFMs) are emerging as reusable representation frameworks for data-driven retrieval, prediction and process modelling within ecohydrology, which integrate EO, meteorological forcing and process models to characterise coupled water, energy and carbon dynamics in vegetation and soil across scales.

By Yi Yu, Jian Peng, Yucheng Lin, Trevor F. Keenan, Thomas F. A. Bishop
arXiv Machine Learning
4d ago

Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface

arXiv:2608. 16134v1 Announce Type: new Abstract: In clinical motor imagery brain-computer interface (MI-BCI) decoding, cross-day transferability and online operation remain two critical challenges.

By Siqi Li (Peking University, Chinese Institute for Brain Research, Beijing), Zhi Li (NeuCyber Neurotech), Tong Liu (NeuCyber Neurotech), Shuai Zhang (NeuCyber Neurotech), Yanfei Jia (Beijing Medical University), Zhiqiang Yi (Beijing Medical University), Jue Xie (NeuCyber Neurotech), Ni Ji (Chinese Academy of Medical Sciences & Peking Union Medical College, Chinese Institute for Brain Research, Beijing)
arXiv Machine Learning
4d ago

Evolving Executable Pipeline Programs for AutoML with Language Models

arXiv:2608. 16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space.

By Sofoklis Kitharidis, Cor J. Veenman, Jan N. van Rijn, Thomas B\"ack, Niki van Stein
arXiv Machine Learning
4d ago

A Low-Cost IoT Device for Environmental Monitoring and Embedded Solar Forecasting with On-Device Incremental Learning

arXiv:2608. 14698v1 Announce Type: cross Abstract: Hyperlocal meteorological sensing is essential for accurate solar photovoltaic forecasting, yet professional-grade meteorological stations require investments easily exceeding 1000~USD per node, making distributed deployments economically inaccessible.

By Erick Michel Lara Pinal, Abhinav Das, Stephan Schl\"uter
arXiv Machine Learning
4d ago

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi