arXiv AI

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

arXiv:2602. 14367v2 Announce Type: replace-cross Abstract: The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation.

arXiv AI
Sep 25

Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

The paper introduces SciMuse, an AI system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. A large-scale evaluation involving over 100 research group leaders across disciplines rated more than 4,400 ideas, yielding modest overall interest scores but showing that 24.9% were rated highly. The study also demonstrates that graph-derived features can predict idea interest and can be used to control idea properties, offering a new methodology for generating and assessing scientific ideas.

By Xuemei Gu, Mario Krenn
arXiv Computation and Language
Aug 27

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.

By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
arXiv AI
Sep 1

Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

Ideation Arena is a battle-style platform that evaluates research ideas generated by large language models (LLMs) and research agents through pairwise human assessment. The system builds shared literature contexts, collects over 6,000 double-blind comparisons from 105 computer science researchers, and constructs an Elo rating leaderboard to rank proposal-stage expert preferences. It also introduces Ideation Arena Eval, a benchmark to test whether automated evaluators align with human preferences, finding that current LLM judges achieve at best 72.56% Soft Accuracy on overall quality.

By Zhiyu Chen, Keyu Zhao, Jigao Fu, Dong Liang, Yanbiao Wu, Jiaoyang Li, Haidong Xue, Xinhua Zeng, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv Computation and Language
Aug 28

DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling

DataSTORM is an LLM‑based agentic system designed to conduct deep research over large‑scale structured databases and internet sources. It applies principles of Exploratory Data Analysis and Data Storytelling to frame research as a thesis‑driven analytical process, iteratively generating hypotheses, performing quantitative reasoning, and crafting coherent narratives. Evaluations on InsightBench and a new ACLED‑based dataset show that DataSTORM surpasses existing systems, achieving significant improvements in insight‑level recall and summary‑level scores.

By Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena, Monica S. Lam