arXiv:2606. 13904v1 Announce Type: cross Abstract: Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions based on intermediate results.
By Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang, Eugene Wu
arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.
By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv:2608. 10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments.
By Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince
arXiv:2607. 27283v1 Announce Type: new Abstract: Long-horizon benchmarks often show that agents fail more as tasks become longer.
By Chao Peng, Zhiheng Lyu, Peijie Dong, Hande Dong, Qiang Lin
arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.
By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
The paper introduces Long-Transduction, a diagnostic framework designed to evaluate how well language models can maintain task fidelity during extended generation tasks that involve continuous reading, mutating, and outputting of context-dependent operations such as arithmetic, sorting, variable lookups, and table transformations. By independently varying local task complexity, input data formatting, and context length, the study isolates failure modes across these axes. Experiments on seven open-weight models reveal significant performance drops—62.8% when scaling context length from 4 to 128K, 36.5% with input format changes, and 39.9% with increased local task complexity—highlighting critical vulnerabilities in long-horizon agentic workflows.
By Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg