arXiv Computation and Language

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research

The paper introduces IDRBench, a framework designed to evaluate how well large language models (LLMs) can integrate knowledge across disciplines for interdisciplinary research. It comprises datasets and tasks—IDR Paper Identification, IDR Idea Integration, and IDR Idea Recommendation—to benchmark LLM performance. The authors analyze ten mainstream LLMs, offering a comprehensive assessment and establishing baselines for future studies.

arXiv AI
Aug 12

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

arXiv:2506. 03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains.

By Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, Xuelong Li
arXiv Machine Learning
Jul 17

Idea2Plan: Exploring AI-Powered Research Planning

arXiv:2510. 24891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant potential to accelerate scientific discovery as valuable tools for analyzing data, generating hypotheses, and supporting innovative approaches in various scientific fields.

By Jin Huang, Silviu Cucerzan, Sujay Kumar Jauhar, Ryen W. White
arXiv AI
Jul 31

Scientific Knowledge Discovery in the Age of Large Language Models

arXiv:2607. 26670v1 Announce Type: cross Abstract: The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depend heavily on manually formulated queries and effortful manual inspection.

By Eleni Adamidi, Serafeim Chatzopoulos, Thanasis Vergoulis
arXiv AI
Jun 12

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

arXiv:2602. 14367v2 Announce Type: replace-cross Abstract: The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation.

By Shuofei Qiao, Yunxiang Wei, Xuehai Wang, Bin Wu, Boyang Xue, Ningyu Zhang, Hossein A. Rahmani, Yanshan Wang, Qiang Zhang, Keyan Ding, Jeff Z. Pan, Huajun Chen, Emine Yilmaz
arXiv AI
Aug 19

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

The paper introduces SGHA, a fully automated system that discovers research problems by structuring scientific literature into evidence-linked objects and a typed evidence graph. SGHA operates entirely on a local 9B open‑weight language model, avoiding proprietary frontier‑model APIs, and outputs traceable research‑problem families with assumptions, objectives, success criteria, and ambiguities. Comparative experiments in five machine‑learning domains show that SGHA’s corpus‑first, evidence‑constrained approach yields inspectable research‑problem formulation without relying on external models.

By Sarvesh Gharat, Junpei Komiyama