arXiv Machine Learning By Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu, Lei Liang, Ningyu Zhang

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

Read the original on arXiv Machine Learning →

LongDS-Bench is a new benchmark for evaluating long-horizon, multi-turn data analysis by agents, featuring 68 tasks derived from real-world Kaggle notebooks that span 2,225 turns across six domains such as Geoscience, Business, and Education. The benchmark focuses on agents’ ability to maintain, update, restore, and compose evolving analytical states, with tasks designed around state-evolution patterns like counterfactual perturbation, rollback, and multi-state composition, and an average dependency span of 11.3 turns. Evaluation of five state-of-the-art models shows that the best model achieves only 48.45% average accuracy, with performance dropping nearly 47 points from early to late turns and long-horizon errors accounting for 52%–69% of failures, indicating that maintaining a correct analytical state is the key bottleneck.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
Aug 12

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv:2608. 10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments.

By Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince
arXiv AI
Jun 29

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.

By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
arXiv AI
3d ago

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

The paper introduces Long-Transduction, a diagnostic framework designed to evaluate how well language models can maintain task fidelity during extended generation tasks that involve continuous reading, mutating, and outputting of context-dependent operations such as arithmetic, sorting, variable lookups, and table transformations. By independently varying local task complexity, input data formatting, and context length, the study isolates failure modes across these axes. Experiments on seven open-weight models reveal significant performance drops—62.8% when scaling context length from 4 to 128K, 36.5% with input format changes, and 39.9% with increased local task complexity—highlighting critical vulnerabilities in long-horizon agentic workflows.

By Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg