Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,813 stories · RSS feed

arXiv AI
Jul 28

Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows

arXiv:2607. 22280v2 Announce Type: replace-cross Abstract: Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, where strong nonlinear interactions produce complex interface deformation, mixing, and multiscale dynamics.

By Harish Ramachandran, Bj\"orn Kimpel, Thomas Paula, Josef Winter, Steffen Schmidt, Nikolaus Adams
arXiv AI
Jul 28

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

arXiv:2607. 23794v1 Announce Type: cross Abstract: Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification.

By Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu
arXiv AI
Jul 28

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

arXiv:2607. 22586v1 Announce Type: new Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens.

By Jinsong Shu, Chenyang Wu, Zhongle Xie, Baokun Wang, Lidan Shou