arXiv AI By Miao Li, Alexander Gurung, Irina Saparina, Mirella Lapata

SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles

Read the original on arXiv AI →

arXiv:2509. 21028v4 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

SciDocBench is a workflow-centered benchmark for scientific document understanding that includes 124 expert-authored questions across seven capability groups and 19 subtasks in five scientific domains. Each question is evaluated under four conditions—English or Chinese, all-images-first or interleaved document representations—resulting in 496 evaluation instances. The benchmark is paired with SciDocIR, a typed evidence-graph representation, and SciDocDataset, a collection of 15K fine-tuning and 8K reinforcement-learning samples, forming an evaluation-to-training framework for scientific-document assistants.

By Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin
arXiv AI
Sep 24

Large Knowledge Model: From Papers to a Scientific Reasoning Landscape

The paper introduces the Large Knowledge Model (LKM), a scientific knowledge infrastructure that converts research literature into shared, computationally accessible reasoning graphs. LKM aligns questions, claims, and reasoning chains across papers, creating a Scientific Reasoning Landscape with Question, Workflow, and Evidence views. The system enhances scientific search, evidence‑grounded QA, and research planning, achieving notable accuracy gains on ChemBench, PubMedQA, and SciBench.

By Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E
arXiv Machine Learning
Sep 21

Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.

By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach