arXiv AI

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

arXiv:2608. 08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines.

arXiv AI
Jul 24

AREX: Towards a Recursively Self-Improving Agent for Deep Research

arXiv:2607. 21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints.

By Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
arXiv AI
Aug 28

DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping

DeepPlanner is an end-to-end reinforcement learning framework designed to enhance the planning capabilities of deep research agents. It introduces an entropy-based advantage shaping mechanism that allocates larger updates to high-entropy planning tokens and selectively upweights sample-level advantages during planning-intensive rollouts. Experiments on seven deep research benchmarks show that DeepPlanner improves planning quality and achieves state‑of‑the‑art results with a lower training budget.

By Wei Fan, Wenlin Yao, Zheng Li, Feng Yao, Xin Liu, Liang Qiu, Qingyu Yin, Yangqiu Song, Bing Yin
arXiv AI
Sep 25

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.

By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung
arXiv AI
Jun 9

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

arXiv:2606. 09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite.

By Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
arXiv AI
4d ago

LongCat-DeepResearch Technical Report

LongCat-DeepResearch is a deep research system that merges an enhanced LongCat model with a multi‑agent workflow to produce comprehensive, evidence‑grounded reports. The workflow separates global planning from detailed investigation, using planning agents to create a ResearchSpec and research agents to draft sections in parallel, followed by targeted local revisions guided by global review. The system achieves strong benchmark scores, including 55.25 on DeepResearchBench and 79.83 on ResearchRubrics, and shows benefits from combining planning perspectives and additional editing for readability.

By Meituan LongCat Team, He Zhu, Yue Xu, Wanli Wu, Haolin Ren, Yuxin Bian, Jiarui Zhao, Rongzhi Zhang, Quanchi Weng, Jinghao Cui, Yu Fan, Yuhan Liu, Yunhu Ye, Jiyuan Ren, Fengcheng Yuan, Zhao Yang, Jiacheng Zhang, Yuchuan Dai, Ruixuan Xiao, Haozhe Sun, Xiangyuan Liu, Cheng Sun, Yao Du, Yiming Hao, Hongbo Guo, Shuo He, Lei Wang, Xunliang Cai, Yan Chen, Fan Yang, Lingchuan Liu
arXiv Machine Learning
Sep 22

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

The paper introduces Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning of large language models cumulative. In the first stage, a multi-agent pipeline uses Monte Carlo Tree Search to explore training strategies while a Distillation Agent records task-specific insights and cross-task confidence scores into a structured repository. In the second stage, SAGE retrieves relevant experience from this repository to guide training on new tasks, achieving a 12.4‑percentage‑point improvement over a baseline pipeline without accumulated experience on nine unseen tasks.

By Haoran Zhao, Wei Du, Dingwen Yang, Jixuan Huang, Junlin Shang, Lingyong Fang, Ya Guo, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv Computation and Language
Aug 27

Mind2Report: Expert-Level Commercial Report Synthesis via Cognitive Deep Research Agent

Mind2Report is a cognitive deep research agent designed to produce expert-level commercial reports from large, noisy web sources. It first clarifies detailed commercial intent to build a structured outline, then recursively gathers and validates evidence into a research memory that evolves with the outline, enabling iterative synthesis of comprehensive reports. The authors also introduce QRC‑Eval, a benchmark of 200 real-world commercial tasks, and show through extensive experiments that Mind2Report outperforms existing proprietary and open-source deep research agents, with ablation studies confirming the contribution of each component.

By Mingyue Cheng, Daoyu Wang, Qi Liu, Shuo Yu, Xiaoyu Tao, Yuqian Wang, Chengzhong Chu, Yu Duan, Mingkang Long, Enhong Chen
arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun