Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

13,227 stories · RSS feed

arXiv Machine Learning
Aug 18

Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment

arXiv:2608. 16005v1 Announce Type: new Abstract: This paper studies the problem of molecule-text alignment, which aims to project molecules and their textual descriptions into a joint latent space for downstream tasks including molecule search and molecular property prediction.

By Shunshun Gu, Shengqi Qiu, Hang Zhou, Xiao Luo
arXiv AI
Aug 18

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

arXiv:2608. 15242v1 Announce Type: new Abstract: When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory.

By Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogang Xie, Dan Pei
arXiv AI
Aug 18

ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search

ATLAS is an embedding‑guided quality‑diversity framework that enables scaffold‑free synthesis of full algorithms for combinatorial optimization using large language models. It allows the LLM to freely choose, restructure, and control algorithm components while automatically detecting and repairing execution, interface, and feasibility failures. Across four NP‑hard problems, ATLAS outperforms state‑of‑the‑art component‑synthesis methods and remains competitive with strong human‑designed algorithms, demonstrating that a larger design space can be practically searched.

By Danial Yazdani, Mohammad Nabi Omidvar, Yuan Sun, Maksud Ibrahimov, Xiaodong Li
arXiv Machine Learning
Aug 18

EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction

EquiPocket is an E(3)-equivariant Graph Neural Network designed to predict ligand binding sites on proteins. It processes proteins as geometric graphs, extracting local surface atom geometry, modeling chemical and spatial relationships, and performing equivariant message passing to capture surface geometry. A dense attention output layer mitigates issues caused by variable protein sizes, and experiments show the method outperforms current state‑of‑the‑art approaches.

By Yang Zhang, Zhewei Wei, Ye Yuan, Chongxuan Li, Wenbing Huang
arXiv AI
Aug 18

GRIP: Grounded Reasoning via Information-Restricted Premises

GRIP (Grounded Reasoning via Information-Restricted Premises) addresses the query dominance problem in retrieval-augmented generation by enforcing a capacity asymmetry: the decoder retains full access to the query while retrieved evidence is funneled through a severe stochastic bottleneck. This design forces the evidence channel to encode only residual information not present in the query. On five reasoning benchmarks, GRIP surpasses strong iterative baselines, reduces query–latent mutual information by about 30×, cuts hallucination by 73%, and its bottleneck outputs occupy subspaces less aligned with the query than baseline representations.

By Lirui Teng
arXiv AI
Aug 18

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

The paper introduces RUPA, a trajectory‑level uncertainty quantification framework for large language model agents. RUPA models an agent’s execution as a directed graph of reasoning states, tool interactions, and environment feedback, then propagates uncertainty across this graph to capture long‑range dependencies. Experiments on benchmarks such as τ‑2, Terminal‑Bench‑2, and GAIA show that RUPA outperforms existing methods, enabling earlier failure detection and more reliable agent execution.

By Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv Machine Learning
Aug 18

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

The paper introduces AMPLIFAI, the first public dataset of multiphase abdominal CT scans annotated with LI-RADS categories and segmented for three key LI-RADS features: arterial phase hyperenhancement, washout, and enhancing capsule. It outlines the dataset’s composition, curation process, and annotation pipeline following the Datasheets for Datasets format to promote transparency and reproducibility. The dataset aims to support the development of AI models for automated hepatocellular carcinoma diagnosis using the biopsy‑free, imaging‑based LI‑RADs framework.

By Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi, Jana Delfino, James Tonascia, Jade Wong-You-Cheong, Barton Lane, Joseph Chirico, Jeffrey D. Hirsch, Ang Li, Heng Huang, Florence X. Doo