arXiv AI

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

arXiv:2606. 18191v1 Announce Type: new Abstract: Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries.

arXiv AI
Sep 25

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.

By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung
arXiv Computer Vision
Aug 28

OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks

OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.

By Jing Wu, Wenjie Ai, Daphne Barretto, Yiye Chen, Qingyu Chen, Yuhang He, Pranit Chawla, Nicholas Gyd\'e, Yanan Jian, Vibhav Vineet
arXiv Computation and Language
Sep 21

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

arXiv:2609.21187v1 Announce Type: new Abstract: Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is...

By Md Tahmid Rahman Laskar, Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong, Tania Habib, Julien Bouvier Tremblay, Shayna Gardiner, Harsh Saini, Matthias Lee, Elena Khasanova, Quinten McNamara, Shashi Bhushan TN
arXiv AI
2d ago

Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization

The paper introduces InFlowOp, a label‑free optimization framework that assigns costs to each decision in a multi‑agent workflow, balancing agent competence against execution time. It determines task granularity and agent assignment before execution and corrects faults during execution using the same cost metric. The authors also present Braid, a benchmark for multi‑agent coordination, and show that InFlowOp outperforms single‑agent baselines by up to 11.97% across various domains.

By Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng, Qingyun Wang, Haifeng Chen
arXiv AI
Aug 21

Inducing Task Models from Computer-Use Traces

arXiv:2608. 20319v1 Announce Type: cross Abstract: Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done.

By Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen, Diyi Yang
arXiv AI
Aug 6

ContextWeave: A Real-World Workflow Benchmark

arXiv:2608. 04830v1 Announce Type: new Abstract: Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.

By Bo Wang, Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu, Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang, Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu
Hugging Face Trending Papers
Aug 5

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams.