arXiv AI By Honglin Bao, Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, James A. Evans

Contemporary AI lacks the imagination to diverge or negate in science

Read the original on arXiv AI →

arXiv:2606. 08251v1 Announce Type: cross Abstract: Bold projections that artificial intelligence will accelerate scientific discovery have raced ahead of evidence from working scientists, and the field still lacks large-scale, scientist-in-the-loop tests of these claims.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

The paper investigates whether large language models (LLMs) can reliably assess scientific hypotheses by using a logit-based energy scoring method that leverages the model’s intrinsic confidence. Across 1,323 papers in 12 disciplines, this intrinsic scoring achieved a 33.0% Hit@1 rate, outperforming a prompted listwise ranking approach that scored 16.6%. The best result, a 1‑billion‑parameter model with logit-based energy scoring, reached 53.1% Hit@1, suggesting that confidence‑based evaluation could improve trustworthy AI‑enabled scientific discovery.

By Swati Rajwal, Sanjay Das, Tirthankar Ghosal
arXiv Computation and Language
Aug 27

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.

By Tim Schopf, Tobias Schreieder, Akiko Aizawa
arXiv AI
Sep 18

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

The paper introduces ScientistTwo, a fully autonomous multi‑agent framework that takes a scientific problem, establishes baselines, generates hypotheses, and coordinates specialized agents to conduct an end‑to‑end discovery cycle without human intervention. It rigorously tests and refines its methods through automated experiments, ablation studies, and a closed‑loop peer‑review engine. Benchmarking against top conferences (ICLR, ICML, NeurIPS) shows that ScientistTwo produces expert‑level, publishable papers and codebases that outperform human state‑of‑the‑art models and receive higher review ratings under automated AI review.

By Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister
arXiv AI
2d ago

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.

By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang