arXiv:2608. 03468v1 Announce Type: new Abstract: Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage.
By Xiuhui You, Jiayi Luo, Zichao Shen, Qingyun Sun, Ziwei Zhang
Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets.
DAGent introduces an Evaluate‑then‑Grow planning approach for deep research agents, building directed acyclic graphs incrementally based on confidence and uncertainty from completed tasks. The framework includes a hierarchical context layer for efficient query handling and a structural reinforcement learning component, DAGRPO, that rewards topology‑conditioned execution. Experiments on BrowseComp‑Plus, GAIA, and xbench‑DeepSearch show DAGent outperforming strong baselines across multiple backbones and scaling to large language models.
By Hanwen Liu, Yuanfu Sun, Qiaoyu Tan
arXiv:2609.35811v1 Announce Type: cross
Abstract: Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent...
By Zongze Wu, Yani Guo, Runnan Li
arXiv:2608. 05225v1 Announce Type: new Abstract: Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic.
By Huirui Xu, Runtao Xu, Shuo Ren, Jiajun Zhang
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
By Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichart