Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

15,794 stories · RSS feed

arXiv Machine Learning
Jul 31

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

arXiv:2607. 27421v1 Announce Type: cross Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints.

By Parishruthi Ganesh, Gerry Dozier, Cheryl Seals
arXiv Machine Learning
Jul 31

Learning to Trace Seiberg Dualities

arXiv:2607. 28628v1 Announce Type: cross Abstract: Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems.

By Jonathan J. Heckman, Shani Meynet, Alessandro Mininno, Gary Shiu
arXiv AI
Jul 31

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

arXiv:2607. 26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied.

By Kawai Chung, Chunkit Chan, Yauwai Yim, Yuxuan Liu, Haochen Shi, Weiqi Wang, Qing Zong, Tianshi Zheng, Yixuan Fu, Kai Chung Wong, Hao Liang, Yifan Gao, Xi Yang, Janet Hui-wen Hsiao, Yangqiu Song
arXiv AI
Jul 31

Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?

arXiv:2607. 12462v2 Announce Type: replace Abstract: Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each node and all nodes across the traffic network.

By Qihang Zhang, Siyao Zhang, Letao Kang, Wenzhe Liang, Miao Zhang, Zhao Zhang