Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,528 stories · RSS feed

arXiv Machine Learning
Jul 31

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

arXiv:2607. 27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions.

By Aakash Bhagat, Shashank Choudhary
arXiv Machine Learning
Jul 31

LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

arXiv:2607. 27704v1 Announce Type: cross Abstract: As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical.

By Sangjin Kim, Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo
arXiv Machine Learning
Jul 31

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

arXiv:2607. 27763v1 Announce Type: cross Abstract: We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions.

By Bowen Wang, Youwen Zhang, Ritesh Mehta
arXiv Machine Learning
Jul 31

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

arXiv:2607. 27945v1 Announce Type: cross Abstract: Sequence models must decide what to write into memory and what to retain.

By Kuo-Chung Peng, Samuel Yen-Chi Chen, Jiun-Cheng Jiang, Chen-Yu Liu, En-Jui Kuo, Yun-Yuan Wang, Tzung-Chi Huang, Prayag Tiwari, Chi-Sheng Chen, Chun-Hua Lin, Yu-Chao Hsu, Tai-Yue Li, Saif Al-Kuwari, Simon See, Kuan-Cheng Chen, Nan-Yow Chen, Hsi-Sheng Goan
arXiv Machine Learning
Jul 31

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

arXiv:2607. 28074v1 Announce Type: cross Abstract: Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset.

By Yash Pandya, Sahil Gupta, Sarthak Harne, Archana Yadav, Kavyansh Chourasia, Hussein Mozannar, Vibhav Vineet, Sara Abdali, Corby Rosset, Yash Lara, Ahmed Awadallah, Ece Kamar, Akshay Nambi
arXiv Machine Learning
Jul 31

Bridging AI and Energy Forecasting: An Autonomous Workflow with Customized Toolkit

arXiv:2307. 07191v3 Announce Type: replace Abstract: Energy forecasting is crucial for the power grid, but fundamentally different from general time series analysis: it highly relies on covariates like meteorological factors, and its goals must align with actual power grid operations, such as risk assessment and system reliability.

By Zhixian Wang, Leandro Von Krannichfeldt, Qingsong Wen, Chaoli Zhang, Liang Sun, Shirui Pan, Yi Wang
arXiv Machine Learning
Jul 31

Region-adaptable retrieval of coastal biogeochemical parameters from near-surface hyperspectral remote sensing reflectance using physics-aware meta-learning

arXiv:2605. 05623v2 Announce Type: replace Abstract: Hyperspectral in situ sensing has shown promise in retrieving aquatic biogeochemical (BGC) parameters, such as total suspended solids, dissolved organic carbon, and total chlorophyll-a, for cost-effective monitoring of coastal water quality.

By Yiqing Guo, Nagur R. C. Cherukuru, Eric A. Lehmann, S. L. Kesav Unnithan, Tim J. Malthus, Gemma Kerrisk, Xiubin Qi, Faisal Islam, Tisham Dhar, Mark J. Doubell
arXiv Machine Learning
Jul 31

ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate

arXiv:2607. 03126v3 Announce Type: replace Abstract: Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their unequal contributions to the reasoning process.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Quan Chen, Yanhua Cheng, Peng Jiang, Binbin Zheng, Xiaolong Liu, Zeyu Chen, Yadong Mu
arXiv Machine Learning
Jul 31

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

arXiv:2508. 04227v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot generalization.

By Yuyang Liu, Qiuhe Hong, Linlan Huang, Alexandra Gomez-Villa, Dipam Goswami, Tiantian Peng, Xialei Liu, Joost van de Weijer, Yonghong Tian