Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
arXiv:2606. 11926v1 Announce Type: cross Abstract: Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction.
Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts.
arXiv:2606. 11926v1 Announce Type: cross Abstract: Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction.
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-end...
The paper introduces Agentic Reasoning for Tree Search (ARTS), a method that uses a reasoning language model to navigate the hypothesis‑experiment space in scientific discovery. Unlike traditional approaches that conflate hypothesis quality with execution quality and prune search logs, ARTS evaluates prior execution logs to distinguish implementation failures from poor hypotheses and selects the next hypothesis to pursue. By employing test‑time training to embed search‑tree knowledge into model weights, ARTS achieves a 15.3% relative improvement over leading algorithms on 22 benchmark tasks and enables smaller models like Qwen3‑4B to match or exceed the performance of larger closed‑source models at lower inference cost.
PrimeScientist is a system that jointly selects research directions and allocates resources for autonomous research agents. It models the problem as a sequential decision task, using an executable plan tree to track competing plans and an adaptive MCTS-based policy to balance exploration and exploitation based on remaining resources and experimental feedback. Experiments on AI research, systems, code optimization, and machine learning engineering show that PrimeScientist improves average reward by 10.3% while reducing research attempts by 50.6% compared to AutoResearch under the same budget.
arXiv:2608. 14354v1 Announce Type: new Abstract: Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration decisions, and computational resources.
arXiv:2609.35561v2 Announce Type: replace Abstract: Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous mo...
arXiv:2606. 07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.
arXiv:2608.31076v1 Announce Type: cross Abstract: Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experi...
Deep search requires agents to answer complex questions through multi-step web search, browsing, evidence comparison, and synthesis. A central challenge is deciding how to search when several directions look plausible but only some will later lead to reliable evidence.
AutoResearch is a two‑stage autonomous research system that links Idea Generation with Idea Execution. In the generation phase it blends new research signals with existing domain knowledge, identifies transferable mechanistic insights, and produces grounded, testable research plans through multi‑model generation and cross‑review. The execution phase then decomposes these plans into experiments, iteratively implements and diagnoses them, and uses independent evidence‑based review to accept or revise conclusions, thereby turning ideas into measurable progress while minimizing hallucinations.
The paper introduces AIDE^2, an AI research agent that recursively improves its own code by proposing, benchmarking, and selecting modifications. Over an eight‑day autonomous run, it achieved seven successive improvements—including new search policies and memory mechanisms—that transferred to four held‑out benchmarks in machine learning, algorithm engineering, and weather forecasting. The agent’s best version matched or outperformed a top human‑engineered production research agent and also reduced reward‑hacking rates, despite never optimizing for that metric.
arXiv:2609.07611v1 Announce Type: new Abstract: Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Exi...