arXiv AI By Gabriel Manso, Emma Fu, Neil Thompson

The AI-Enabled Scientific Frontier

Read the original on arXiv AI →

The paper "The AI-Enabled Scientific Frontier" analyzes 2,507 head‑to‑head comparisons of AI versus other scientific methods across 27 disciplines from 2000 to early 2025. It finds that AI often outperforms traditional statistics but at higher computational cost, while in about a quarter of cases AI is both more expensive and less effective. Compared to scientific computing, AI usually underperforms but at lower cost, though since 2020 its performance has improved, now surpassing computing in over half of the comparisons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

AI in Science: Early Insights

The paper "AI in Science: Early Insights" analyzes AI’s impact on scientific work using data from 15 million Gemini interactions, 2,600 specialized AI models, and a survey of 600 scientists. It finds widespread AI adoption, complementary use of large language models and specialized tools, significant productivity gains of about seven hours per week, and a shift in research bottlenecks toward hypothesis backlog and verification needs. The study suggests AI can boost scientific productivity but its full effect depends on addressing new downstream challenges.

By Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri, Arthur Turrell, Julian Jacobs, Atoosa Kasirzadeh, Ana Trisovic, Yiyuan Chen, Tanya Rodchenko, Catherine Pollard, Scott Strand, Daniel Rock, Zanna Iscenko, Fabien Curto Millet, Neil Thompson, James Manyika
arXiv AI
Jun 16

Artificial Intelligence Index Report 2026

arXiv:2606. 15708v1 Announce Type: new Abstract: Welcome to the ninth edition of the AI Index report.

By Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, Dan Weld
arXiv AI
2d ago

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.

By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister