Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,813 stories · RSS feed

arXiv AI
Jul 28

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

arXiv:2607. 22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable.

By Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du
arXiv AI
Jul 28

An Ontology for Machine Learning Interatomic Potentials

arXiv:2607. 23219v1 Announce Type: new Abstract: Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost.

By Daniel Hern\'andez, Jong Hyun Jung, Yuji Ikeda, Yongliang Ou, Pranav Kumar, Tom Sch\"achtel, Wenchuan Liu, Xin Li, Xi Zhang, Xiang Xu, Lifang Zhu, Fritz K\"ormann, Steffen Staab, Blazej Grabowski
arXiv AI
Jul 28

Do LLMs Know Their Vulnerable Scenarios?

arXiv:2607. 23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.

By Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
arXiv AI
Jul 28

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

arXiv:2607. 24280v1 Announce Type: new Abstract: Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision.

By Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou
arXiv AI
Jul 28

Comparing Optimization Models for Radiotherapy Scheduling

arXiv:2607. 22539v1 Announce Type: cross Abstract: The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task that has a massive impact on clinical outcomes given the central role of radiotherapy in cancer care.

By C. C. Rambaldi Migliore, D. Stanicel, N. Musliu, G. Iacca, M. Roveri
arXiv AI
Jul 28

Real-time Reconstruction of Human Visual Perception from fMRI

arXiv:2607. 22753v1 Announce Type: cross Abstract: Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances.

By Rishab S. Iyer, Jiaxin Cindy Tu, Cesar Kadir Torrico Villanueva, Anish Mahishi, Ross P. Kempner, Jacob S. Prince, Ernest W. Lo, Akash Bhowmick, Hritik Arasu, Amaar Chughtai, Elizabeth A. McDevitt, Paul S. Scotti, Kenneth A. Norman
arXiv AI
Jul 28

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

arXiv:2607. 23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers.

By Xinyan Zhong, Yuwei Shi, Yuqi Wei, Chen Shen, Tianhang Zhou, Zhenghao Wu