Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

22,245 stories · RSS feed

arXiv AI
Jun 10

Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix

arXiv:2606. 09884v1 Announce Type: cross Abstract: We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high event rates.

By Shree Murthy, Rohan Pandey
arXiv Machine Learning
Jun 10

Encoding the Euler Characteristic Transform

arXiv:2606. 10824v1 Announce Type: new Abstract: The Euler Characteristic Curve (ECC) records the Euler characteristic of a linearly embedded cell complex as a function of filtration height in a given direction, and the Euler Characteristic Transform (ECT) is the injective shape descriptor obtained by collecting ECCs over many directions.

By Nello Blaser, Odin Hoff Gardaa, Lars M. Salbu, Elena Xinyi Wang, Bastian Rieck
arXiv Machine Learning
Jun 10

Learning Doubly Sparse Explicitly Conditioned Transforms

arXiv:2606. 10975v1 Announce Type: new Abstract: Finding convenient spaces in which certain hypotheses regarding an assumed sparse structure of natural signals hold true has become a desirable result in recent research, its implications being reflected in areas such as data compression, noise reduction and feature extraction.

By Tudor Pistol
arXiv Machine Learning
Jun 10

Inverse Probability Weighting and Age-of-Information Aggregation for Decentralized Federated Learning under Partial Reception

arXiv:2606. 10774v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) over lossy wireless networks faces two key challenges: selection bias, where updates from poor-quality links are systematically underrepresented due to partial model reception, and update staleness, where asynchronous nodes contribute outdated information.

By Chanuka A. S. Hewa Kaluannakkage, Rajkumar Buyya
arXiv AI
Jun 10

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

arXiv:2606. 10917v1 Announce Type: new Abstract: Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static training environments, which hinder broader generalization.

By Xucong Wang, Ziyu Ma, Shidong Yang, Tongwen Huang, Pengkun Wang, Yong Wang, Xiangxiang Chu