arXiv Computation and Language By Yuanhao Shen, Daniel Xavier de Sousa, Ricardo Mar\c{c}al, Hongyu Guo, Xiaodan Zhu

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research

Read the original on arXiv Computation and Language →

The paper introduces IDRBench, a framework designed to evaluate how well large language models (LLMs) can integrate knowledge across disciplines for interdisciplinary research. It comprises datasets and tasks—IDR Paper Identification, IDR Idea Integration, and IDR Idea Recommendation—to benchmark LLM performance. The authors analyze ten mainstream LLMs, offering a comprehensive assessment and establishing baselines for future studies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 12

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

arXiv:2506. 03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains.

By Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, Xuelong Li