The paper investigates whether large language models (LLMs) can reliably assess scientific hypotheses by using a logit-based energy scoring method that leverages the model’s intrinsic confidence. Across 1,323 papers in 12 disciplines, this intrinsic scoring achieved a 33.0% Hit@1 rate, outperforming a prompted listwise ranking approach that scored 16.6%. The best result, a 1‑billion‑parameter model with logit-based energy scoring, reached 53.1% Hit@1, suggesting that confidence‑based evaluation could improve trustworthy AI‑enabled scientific discovery.
By Swati Rajwal, Sanjay Das, Tirthankar Ghosal
arXiv:2606. 12071v1 Announce Type: cross Abstract: LLMs are increasingly used to generate and judge scientific ideas.
By Soumitra Sinhahajari, Navonil Majumder, Soujanya Poria
Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows.
The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.
By Tim Schopf, Tobias Schreieder, Akiko Aizawa
The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.
By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.
By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang
The paper introduces SciMuse, an AI system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. A large-scale evaluation involving over 100 research group leaders across disciplines rated more than 4,400 ideas, yielding modest overall interest scores but showing that 24.9% were rated highly. The study also demonstrates that graph-derived features can predict idea interest and can be used to control idea properties, offering a new methodology for generating and assessing scientific ideas.
By Xuemei Gu, Mario Krenn
arXiv:2607. 28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery.
By Vaibhava Lakshmi Ravideshik, Mayank Kejriwal
arXiv:2606. 25198v2 Announce Type: replace Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning.
By Antonis Antoniades, Deepak Nathani, Ritam Saha, Alfonso Amayuelas, Ivan Bercovich, Zhaotian Weng, Vignesh Baskaran, Kunal Bhatia, William Yang Wang
The paper "AI in Science: Early Insights" analyzes AI’s impact on scientific work using data from 15 million Gemini interactions, 2,600 specialized AI models, and a survey of 600 scientists. It finds widespread AI adoption, complementary use of large language models and specialized tools, significant productivity gains of about seven hours per week, and a shift in research bottlenecks toward hypothesis backlog and verification needs. The study suggests AI can boost scientific productivity but its full effect depends on addressing new downstream challenges.
By Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri, Arthur Turrell, Julian Jacobs, Atoosa Kasirzadeh, Ana Trisovic, Yiyuan Chen, Tanya Rodchenko, Catherine Pollard, Scott Strand, Daniel Rock, Zanna Iscenko, Fabien Curto Millet, Neil Thompson, James Manyika
arXiv:2609.07611v1 Announce Type: new
Abstract: Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Exi...
By Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See
arXiv:2605. 10574v3 Announce Type: replace Abstract: As artificial intelligence advances, models are not improving uniformly.
By Shray Mathur, J. Anibal Boscoboinik, Esther H. R. Tsai, Kevin G. Yager