arXiv Machine Learning By Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Read the original on arXiv Machine Learning →

arXiv:2608. 00355v1 Announce Type: cross Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

The paper investigates why some tasks in the Terminal‑Bench/Frontier‑Bench datasets fail for all agents, distinguishing genuine difficulty from artifacts such as missing context, broken solutions, infrastructure failures, or verifier bypasses. Analyzing 125 all‑fail tasks, only 78 are certified as genuinely unsolved after applying a validity screen; the rest are attributable to broken oracles, infrastructure issues, bypassable verifiers, or insufficient evidence. The study concludes that a zero pass rate does not automatically indicate a hard task and recommends that frontier benchmarks provide evidence for all‑fail tasks before claiming capability gaps.

By Edward Lue Chee Lip, Boden Moraski, Tim Knappe, Lang Xiong, Sarvesh Gharat, Antonio Mari, Ivan Bercovich
arXiv AI
Sep 3

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

The paper investigates why large language model (LLM) agents fail on long, multi‑step production workflows despite high benchmark success. By testing nine models (1.2 B–671 B parameters) across six task families and multiple horizons, the authors find that task success follows a geometric decay governed by a per‑step reliability that never reaches 1, leading to inevitable collapse for long horizons. The degradation is driven mainly by step count rather than context length, and the study quantifies a significant gap between benchmark and production performance, especially for agentic tool‑use tasks.

By Shubhra Mittal
arXiv AI
2d ago

It Takes Workflows to Evolve Better Workflows

arXiv:2610.01026v1 Announce Type: cross Abstract: Tackling complex real-world tasks can exceed the capabilities of a single large language model (LLM), motivating the use of multi-agent workflows tha...

By Xuehang Guo, Haoyu Wang, Haifeng Chen, Yangyi Chen, Zhenhailong Wang, Qingyun Wang