arXiv AI By Huirui Xu, Runtao Xu, Shuo Ren, Jiajun Zhang

Project2Task: Graph-Guided Project-Level Planning for Autonomous Research

Read the original on arXiv AI →

arXiv:2608. 05225v1 Announce Type: new Abstract: Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

DAGent introduces an Evaluate‑then‑Grow planning approach for deep research agents, building directed acyclic graphs incrementally based on confidence and uncertainty from completed tasks. The framework includes a hierarchical context layer for efficient query handling and a structural reinforcement learning component, DAGRPO, that rewards topology‑conditioned execution. Experiments on BrowseComp‑Plus, GAIA, and xbench‑DeepSearch show DAGent outperforming strong baselines across multiple backbones and scaling to large language models.

By Hanwen Liu, Yuanfu Sun, Qiaoyu Tan
arXiv AI
Sep 17

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

PrimeScientist is a system that jointly selects research directions and allocates resources for autonomous research agents. It models the problem as a sequential decision task, using an executable plan tree to track competing plans and an adaptive MCTS-based policy to balance exploration and exploitation based on remaining resources and experimental feedback. Experiments on AI research, systems, code optimization, and machine learning engineering show that PrimeScientist improves average reward by 10.3% while reducing research attempts by 50.6% compared to AutoResearch under the same budget.

By Xinle Yu, Fan Bai, Kaiser Sun, Hengshuo Miao, Abhay Anand, Zhongyan Luo, Kun Zhou, Zhen Wang
Hugging Face Trending Papers
Jul 29

SciDataSailor: Deep Scientific Data Exploring

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.