Measuring the Creativity of Frontier LLMs in Automated Research
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 01233v1 Announce Type: cross Abstract: LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference.
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
arXiv:2606. 12071v1 Announce Type: cross Abstract: LLMs are increasingly used to generate and judge scientific ideas.
arXiv:2608.30047v1 Announce Type: new Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they...
The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.