Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

17,750 stories · RSS feed

arXiv Machine Learning
Jul 7

Local Constrained Bayesian Optimization

arXiv:2603. 07965v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) for high-dimensional constrained problems remains a significant challenge due to the curse of dimensionality.

By Jing Jingzhe, Fan Zheyi, Szu Hui Ng, Qingpei Hu
arXiv AI
Jul 7

Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration

arXiv:2511. 02200v2 Announce Type: replace Abstract: The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling diverse agents to integrate unique expertise, collaborate flexibly, and address challenges unattainable for individual models.

By Jingbo Wang, Sendong Zhao, Haochun Wang, Yuzheng Fan, Ting Liu
arXiv Machine Learning
Jul 7

Graph Neural Networks for the Graphical Bootstrap

arXiv:2607. 03109v1 Announce Type: cross Abstract: We study a graph classification problem involving over 20 million graphs, arising from high-order perturbative computations of correlators in planar $\mathcal{N}=4$ super-Yang--Mills, a model closely related to the theory of the strong nuclear force.

By Rigers Aliaj, Gabriele Dian, Reza Doobary, Paul Heslop
arXiv AI
Jul 7

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

arXiv:2607. 04854v1 Announce Type: new Abstract: Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications.

By Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu