arXiv Computation and Language

Measuring the Creativity of Frontier LLMs in Automated Research

arXiv Computation and Language
Aug 27

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.

By Tim Schopf, Tobias Schreieder, Akiko Aizawa
arXiv Computation and Language
Sep 7

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

The paper proposes a new method for measuring the novelty of biomedical research by calculating latent distances between knowledge units—specifically MeSH terms—using three types of relationships: network, semantic, and hierarchical. It demonstrates that each relationship captures distinct distances and that combining all three yields a more accurate novelty assessment than existing metrics. Validation on a large PLoS ONE dataset and a H1 Connect dataset shows stronger alignment with peer judgments compared to prior indicators.

By Yi Zhao, Heng Zhang, Yuzhuo Wang, Wenqing Wu, Tong Bao, Chengzhi Zhang