Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

17,521 stories · RSS feed

arXiv AI
Jul 7

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.

By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang
arXiv AI
Jul 7

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

arXiv:2607. 03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation.

By Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
arXiv AI
Jul 7

Spectral Signatures of Large Language Models

arXiv:2607. 03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation.

By Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang
arXiv AI
Jul 7

Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation

arXiv:2607. 03487v1 Announce Type: cross Abstract: Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate estimators on simplified, low-dimensional distributions, leaving their performance on complex, realistic data largely unexplored.

By Alberto Foresti, Ivan Butakov, Alexander Tolmachev, Giulio Franzese, Alexey Frolov, Pietro Michiardi
arXiv AI
Jul 7

Graph Unitary Message Passing

arXiv:2403. 11199v2 Announce Type: replace-cross Abstract: Unitarity is a useful principle for stabilizing deep neural networks, but in graph neural networks (GNNs) instability is induced not only by learnable parameters but also by the graph propagation operator.

By Haiquan Qiu, Quanming Yao
arXiv AI
Jul 7

LLMoxie: Exploring Agentic AI for Scientific Software Development

arXiv:2607. 02703v1 Announce Type: cross Abstract: In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a LiteLLM/MLflow control plane for authentication, budgeting, PII masking, and observability, and an application augmentation layer for AI coding agents.

By Landung Setiawan, Anant Mittal, Cordero Core, Anshul Tambay, Carlos Garcia Jurado Suarez, David A. C. Beck, Andrew J. Connolly, Vani Mandava
arXiv AI
Jul 7

HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

arXiv:2607. 04329v1 Announce Type: new Abstract: Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers.

By Yaozu Wu, Wei-Chieh Huang, Jizhou Guo, Dongyuan Li, Renhe Jiang, Henry Peng Zou, Chunyu Miao, Shanghao Li, Weizhi Zhang, WeiWei Ye, Yankai Chen, Meng Zhang, Xue Liu, Philip S. Yu
arXiv AI
Jul 7

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

arXiv:2607. 03748v1 Announce Type: new Abstract: Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge.

By Zican Hu, Xuyang Hu, Yiming Liu, Zuwei Long, Wei Liu, Yunzhuo Hao, Jiawei Gu, Linjie Li, Yu Cheng, Zhenhong Sun, Weibo Gu, Xing Sun, Zhi Wang