NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.
By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2606. 12071v1 Announce Type: cross Abstract: LLMs are increasingly used to generate and judge scientific ideas.
By Soumitra Sinhahajari, Navonil Majumder, Soujanya Poria
arXiv:2603.20884v4 Announce Type: replace
Abstract: To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for pa...
By Jiajun Hou, Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Xiaopeng Ke, Derek F. Wong, Min Zhang
arXiv:2510. 27313v3 Announce Type: replace Abstract: Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging.
By Philipp Davydov, Ameya Prabhu, Matthias Bethge, Elisa Nguyen, Seong Joon Oh
arXiv:2608. 14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result.
By Ayrton Porto
arXiv:2607. 26066v1 Announce Type: cross Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review.
By Ranjitha Shivaprasad Ballakuraya, Arash Mahyari, Ashok Srinivasan
The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.
By Tim Schopf, Tobias Schreieder, Akiko Aizawa
The paper proposes a new method for measuring the novelty of biomedical research by calculating latent distances between knowledge units—specifically MeSH terms—using three types of relationships: network, semantic, and hierarchical. It demonstrates that each relationship captures distinct distances and that combining all three yields a more accurate novelty assessment than existing metrics. Validation on a large PLoS ONE dataset and a H1 Connect dataset shows stronger alignment with peer judgments compared to prior indicators.
By Yi Zhao, Heng Zhang, Yuzhuo Wang, Wenqing Wu, Tong Bao, Chengzhi Zhang
The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.
By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang
Large language models are increasingly used as automated reviewers in scientific evaluation, creating a recursive feedback loop where later reviewers learn from earlier model-generated judgments. A study using Llama 3.1 8B fine‑tuned on ICLR reviews shows that incorporating synthetic reviews compresses rating distributions and reduces semantic diversity, a phenomenon termed scientific‑judgment collapse. To counter this, the authors introduce TrustReviewer, an open‑source LLM system that curates training data and applies paired activation steering at test time to preserve judgment diversity and improve recommendation alignment.
By Sy-Tuyen Ho, Minghui Liu, Furong Huang
arXiv:2606. 25198v2 Announce Type: replace Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning.
By Antonis Antoniades, Deepak Nathani, Ritam Saha, Alfonso Amayuelas, Ivan Bercovich, Zhaotian Weng, Vignesh Baskaran, Kunal Bhatia, William Yang Wang
arXiv:2605. 05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences.
By Nicholas S. Kersting, Vittorio Castelli, Chieh Ting Yeh, Xinzhu Wang, Saad Taame, Khaoula Allak