arXiv:2609.22104v1 Announce Type: new
Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottle...
By Rongcan Pei, Fang Guo, Qinglin Qi, Qi Zhu, Yun Luo, Jianhao Yan, Minjun Zhu, Qiujie Xie, Dehong Zheng, Yue Zhang
arXiv:2608. 13136v1 Announce Type: cross Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention.
By Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen
The paper introduces SciMuse, an AI system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. A large-scale evaluation involving over 100 research group leaders across disciplines rated more than 4,400 ideas, yielding modest overall interest scores but showing that 24.9% were rated highly. The study also demonstrates that graph-derived features can predict idea interest and can be used to control idea properties, offering a new methodology for generating and assessing scientific ideas.
By Xuemei Gu, Mario Krenn
IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.
By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2606. 07226v1 Announce Type: cross Abstract: Human creativity has emerged as a critical competency in the era of large language models.
By Tongzhou Yu, Mingjia Li, Hong Qian, Wenkai Wang, Zongbao Zhang, Yaoyu Jiang, Xiangfeng Wang, Aimin Zhou, Jiajun Guo
arXiv:2607. 04439v1 Announce Type: new Abstract: Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions.
By Qihao Zhao, Yangyu Huang, Yalun Dai, Lingao Xiao, Jianjun Gao, Xin Zhang, Wenshan Wu, Scarlett Li, Yang He, Yan Lu, Yap Kim Hui
arXiv:2604. 09497v2 Announce Type: replace-cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.
By Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Emmanuel Malherbe, C\'eline Hudelot, Pierre Colombo
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
By Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya, Mohor Banerjee, Swaagat Bikash Saikia, Alvin Chan
Ideation Arena is a battle-style platform that evaluates research ideas generated by large language models (LLMs) and research agents through pairwise human assessment. The system builds shared literature contexts, collects over 6,000 double-blind comparisons from 105 computer science researchers, and constructs an Elo rating leaderboard to rank proposal-stage expert preferences. It also introduces Ideation Arena Eval, a benchmark to test whether automated evaluators align with human preferences, finding that current LLM judges achieve at best 72.56% Soft Accuracy on overall quality.
By Zhiyu Chen, Keyu Zhao, Jigao Fu, Dong Liang, Yanbiao Wu, Jiaoyang Li, Haidong Xue, Xinhua Zeng, Yuanyi Zhen, Fengli Xu, Yong Li
DataSTORM is an LLM‑based agentic system designed to conduct deep research over large‑scale structured databases and internet sources. It applies principles of Exploratory Data Analysis and Data Storytelling to frame research as a thesis‑driven analytical process, iteratively generating hypotheses, performing quantitative reasoning, and crafting coherent narratives. Evaluations on InsightBench and a new ACLED‑based dataset show that DataSTORM surpasses existing systems, achieving significant improvements in insight‑level recall and summary‑level scores.
By Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena, Monica S. Lam
arXiv:2605.29522v2 Announce Type: replace
Abstract: As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for bo...
By Ziyue Yang, Da Ma, Hanqi Li, Zijian Wang, Tiancheng Huang, Zijian Hu, Chenrun Wang, Yunzhe Zhang, Xiaobao Wu, Kai Yu, Lu Chen