arXiv Computation and Language

Can Large Language Models Forecast What Researchers Study Next?

arXiv Computation and Language
23h ago

Can Language Models Learn to Forecast Stock Prices

arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...

By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv AI
Jun 9

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

arXiv:2606. 08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence.

By Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou, Wenjing Gong, Kaili Zhang, Bingqian Chen, Mitch Zhang, Yifan Yang
arXiv AI
Jul 10

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

arXiv:2607. 08046v1 Announce Type: cross Abstract: Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast.

By Rapha\"el Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
arXiv AI
Aug 25

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The paper introduces Lit2Test, a benchmark that evaluates language models’ research idea proposals by requiring each idea to include a falsifiable outcome, thereby making quality decidable. Built from 200 real-paper neighborhoods, the benchmark gathers proposals from four frontier models and compares them via 1,200 blind pairwise judgments, with reliability checks and human calibration. The results show a consistent ranking of the models, driven by test and metric quality rather than fluency, and the authors release the benchmark and related artifacts for public use.

By Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Hongyao Zuo (Tianjin University), Ziwen Gong (Hainan University), Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Tong Yang (Peking University), Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Xiaohui Li (Huawei Technologies), Haoli Bai (Huawei Technologies)