Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

22,912 stories · RSS feed

arXiv Machine Learning
Jun 10

Encoding the Euler Characteristic Transform

arXiv:2606. 10824v1 Announce Type: new Abstract: The Euler Characteristic Curve (ECC) records the Euler characteristic of a linearly embedded cell complex as a function of filtration height in a given direction, and the Euler Characteristic Transform (ECT) is the injective shape descriptor obtained by collecting ECCs over many directions.

By Nello Blaser, Odin Hoff Gardaa, Lars M. Salbu, Elena Xinyi Wang, Bastian Rieck
arXiv AI
Jun 10

RAT: Reference-Augmented Training for ASV Anti-Spoofing

arXiv:2606. 10908v1 Announce Type: cross Abstract: We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference.

By Vojt\v{e}ch Stan\v{e}k, Anton Firc, Jakub Re\v{s}, Kamil Malinka
arXiv AI
Jun 10

Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix

arXiv:2606. 09884v1 Announce Type: cross Abstract: We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high event rates.

By Shree Murthy, Rohan Pandey
arXiv AI
Jun 10

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.

By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan