Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,030 stories · RSS feed

arXiv AI
Aug 11

Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs

arXiv:2608. 07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executions are non-deterministic, and AI workloads bring high-dimensional inputs and non-linear operations that defeat fuzzing and symbolic execution in isolation.

By Gautham Koorma, Vikas Sharma, George Edwards, Mahdi Eslamimehr
arXiv AI
Aug 11

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.

By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv AI
Aug 11

Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

arXiv:2608. 08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content.

By Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
arXiv AI
Aug 11

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

arXiv:2512. 05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes.

By Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang
arXiv AI
Aug 11

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

arXiv:2608. 09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit.

By Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen
arXiv Machine Learning
Aug 11

Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation

arXiv:2608. 07562v1 Announce Type: cross Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution.

By Nafis Fuad, Xiaodong Qian, Dongxiao Zhu