arXiv:2606. 19704v1 Announce Type: new Abstract: Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes.
Machine-generated by The Flow from the publisher's headline and feed description
— not written or checked by a human. The full article lives at arXiv AI.
arXiv:2609.07785v1 Announce Type: new
Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that c...
arXiv:2602. 03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget.
Agent evaluations increasingly benchmark LLMs, but rankings can be swayed by evaluation conditions such as scaffolds or tasks, making reliability claim‑dependent. A Bayesian variance‑decomposition framework applied to 22 benchmarks shows that reliability varies with the measurement goal: fixed model‑scaffold systems rank reliably, while underlying‑model rankings are less stable. Scaffold choice can alter conclusions, and adding more tasks only modestly improves reliability when scaffold coverage is limited; however, pooling diverse benchmarks can substantially raise cross‑task ranking reliability and reduce cost.
By Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
arXiv:2607. 25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules.
By Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen
arXiv:2603. 22376v2 Announce Type: replace-cross Abstract: We present an AI Co-Scientist framework that closes the research loop for the production search-ranking system of a large online travel platform -- pairing LLM agents with direct cloud-compute access so that idea generation, code implementation, GPU experimentation, and result analysis iterate end-to-end with a human scientist in the loop.