Introducing deep research
An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you. Available to Pro users today, Plus and Team next.
Related stories
Mind2Report: Expert-Level Commercial Report Synthesis via Cognitive Deep Research Agent
Mind2Report is a cognitive deep research agent designed to produce expert-level commercial reports from large, noisy web sources. It first clarifies detailed commercial intent to build a structured outline, then recursively gathers and validates evidence into a research memory that evolves with the outline, enabling iterative synthesis of comprehensive reports. The authors also introduce QRC‑Eval, a benchmark of 200 real-world commercial tasks, and show through extensive experiments that Mind2Report outperforms existing proprietary and open-source deep research agents, with ablation studies confirming the contribution of each component.
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
arXiv:2606. 13710v1 Announce Type: new Abstract: Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence.
DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents
arXiv:2609.36344v1 Announce Type: new Abstract: Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they m...
S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents
arXiv:2606. 15367v1 Announce Type: new Abstract: Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and report generation.
Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation
arXiv:2606. 05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference.
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
Deep Research Bench II is a new benchmark designed to evaluate Deep Research Agents (DRAs) by requiring them to produce research reports for 132 grounded tasks across 22 domains. Each report is assessed using 9,430 fine‑grained binary rubrics that cover information recall, analysis, and presentation, all derived from expert‑written investigative articles through a rigorous LLM‑plus‑human pipeline. Evaluation of current state‑of‑the‑art DRAs shows that even the best models satisfy fewer than 50% of these rubrics, highlighting a significant gap between automated agents and human experts.
DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation
arXiv:2609.22104v1 Announce Type: new Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottle...
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
arXiv:2606. 07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.
LongCat-DeepResearch Technical Report
LongCat-DeepResearch is a deep research system that merges an enhanced LongCat model with a multi‑agent workflow to produce comprehensive, evidence‑grounded reports. The workflow separates global planning from detailed investigation, using planning agents to create a ResearchSpec and research agents to draft sections in parallel, followed by targeted local revisions guided by global review. The system achieves strong benchmark scores, including 55.25 on DeepResearchBench and 79.83 on ResearchRubrics, and shows benefits from combining planning perspectives and additional editing for readability.
AIM: Agentic Idea Management for Automated Research
arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven se...
Investing in multi-agent AI safety research
Google DeepMind and partners announce a $10M funding call for multi-agent safety research.