Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,257 stories · RSS feed

arXiv AI
Jul 8

MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data

arXiv:2605. 22775v2 Announce Type: replace-cross Abstract: Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handling frequent data missingness from blinks and tracking failures, and efficiently modeling long-range temporal dependencies.

By Amir Mousavi, Mohammad Sadegh Sirjani, Erfan Nourbakhsh, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles
arXiv AI
Jul 8

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv:2607. 05614v1 Announce Type: cross Abstract: Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications.

By Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha, Ezharuddin Jubaer, Armun Alam, Pranjal Kumar Nandi, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman
arXiv AI
Jul 8

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

arXiv:2607. 06309v1 Announce Type: cross Abstract: Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities.

By Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
arXiv Machine Learning
Jul 8

A semantic mutation metric for metamorphic relation adequacy in scientific computing programs

arXiv:2605. 17437v2 Announce Type: replace-cross Abstract: Context.

By Meng Li (School of Computing, University of South China, Hengyang, China, Hunan Engineering Research Center of Software Evaluation and Testing for Intellectual Equipment, Hengyang, China, CNNC Key Laboratory on High Trusted Computing, Hengyang, China), Xiaohua Yang (School of Computing, University of South China, Hengyang, China, Hunan Engineering Research Center of Software Evaluation and Testing for Intellectual Equipment, Hengyang, China, CNNC Key Laboratory on High Trusted Computing, Hengyang, China), Jie Liu (School of Computing, University of South China, Hengyang, China, Hunan Engineering Research Center of Software Evaluation and Testing for Intellectual Equipment, Hengyang, China, CNNC Key Laboratory on High Trusted Computing, Hengyang, China), Shiyu Yan (School of Computing, University of South China, Hengyang, China, Hunan Engineering Research Center of Software Evaluation and Testing for Intellectual Equipment, Hengyang, China, CNNC Key Laboratory on High Trusted Computing, Hengyang, China)
arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera
arXiv Machine Learning
Jul 8

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

arXiv:2602. 19710v3 Announce Type: replace-cross Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision.

By Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu