arXiv AI By Guojun Liao

A Three-Layer Framework for AI in Scientific Discovery

Read the original on arXiv AI →

arXiv:2606. 13566v1 Announce Type: new Abstract: Current discussions of AI in scientific discovery are often dominated by two visible capabilities: search over existing knowledge and execution through optimization, simulation, and automation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

The Problem Is the Problem: Towards Scalable Mathematical Discovery

The paper introduces a new human‑AI collaboration paradigm for mathematical discovery, shifting from selecting individual problems to exploring broad research directions. It presents the Find, Attempt, and Recommend (FAR) pipeline, which automatically searches a literature corpus, filters candidate conjectures, and surfaces promising resolutions for expert review. In a combinatorics pilot, FAR processed over 5,000 papers, identified thousands of open conjectures, and ultimately highlighted 77 items that led to new discoveries.

By Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
arXiv AI
Aug 19

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

The paper introduces SGHA, a fully automated system that discovers research problems by structuring scientific literature into evidence-linked objects and a typed evidence graph. SGHA operates entirely on a local 9B open‑weight language model, avoiding proprietary frontier‑model APIs, and outputs traceable research‑problem families with assumptions, objectives, success criteria, and ambiguities. Comparative experiments in five machine‑learning domains show that SGHA’s corpus‑first, evidence‑constrained approach yields inspectable research‑problem formulation without relying on external models.

By Sarvesh Gharat, Junpei Komiyama
arXiv Machine Learning
Sep 11

Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

The paper introduces HorizonMath, a benchmark of 113 largely unsolved mathematical problems across eight domains, paired with an open-source framework for automated verification. It focuses on the generator‑verifier gap, targeting problems that are hard to discover but easy to verify computationally, thereby avoiding costly formal proof verification or manual review. Using this framework, the authors found six novel solutions—three each from GPT‑5.4 Pro and GPT‑5.6 Sol—demonstrating that current models can contribute to mathematical research, while most state‑of‑the‑art models score below 10%.

By Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen, Eliot Hodges, Dulhan Jayalath, Charles London, Kalyan Ramakrishnan, Jakob Foerster, Cheng Zhang, Flaviu Cipcigan, Philip Torr, Alessandro Abate
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister