arXiv:2609.24246v1 Announce Type: new
Abstract: Large language models (LLMs) have shown remarkable progress in natural language understanding, yet their effectiveness in specialized fields like astro...
By Akhil Sharma, Jatin Gupta, Ali Imam Abidi
The paper investigates whether domain-specific fine‑tuning benefits open‑ended scientific reasoning in astronomy. Using a curated 300‑question QA benchmark from 2017–2026 Olympiad‑style materials, the authors compare open‑weight, API‑served general‑purpose, multimodal, and astronomy‑specialized language models. Results show that strong general‑purpose models set the highest correctness baseline, but variations in metric agreement, judge sensitivity, benchmark composition, and modality suggest that domain specialization is task‑ and deployment‑dependent and that domain‑specific evaluation is crucial for scientific workflows.
By Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang
The paper presents a SciBERT-based method for automatically classifying scientific papers into four telescope-related categories—science, instrumentation, mention, and not telescope—within strict 512-token limits. Despite truncation challenges, the approach achieved a macro F1 score of 0.89, topping the WASP-2025 leaderboard. The authors analyze truncation effects, compare chunking and long-context models, and offer insights into efficient scientific text curation.
By Madhusudhana Naidu
An agentic framework called GW‑Eyes, powered by large language models, is introduced to autonomously associate gravitational‑wave (GW) signals with candidate electromagnetic (EM) counterparts. It integrates domain‑specific tools for tasks such as catalog management, skymap visualization, and rapid verification, while enabling natural‑language interaction to assist human experts. The framework leverages LLMs’ decision‑making and traceable reasoning to address the growing data‑analysis challenges of next‑generation GW and EM detectors.
By Yiming Dong, Yacheng Kang, Junjie Zhao, Xinyuan Zhu, Ziming Wang, Lijing Shao
SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.
By Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai
arXiv:2609.13648v1 Announce Type: new
Abstract: Solar energy decision support is fragmented across dashboards that provide data without explanation, research papers are slow to parse, and general-pur...
By Jyotsna Singh
The paper proposes a new data interpretation stage that transforms spatiotemporal field data into physically meaningful quantities before feeding them to a large language model for partial differential equation (PDE) discovery. On simulated benchmarks, this approach nearly triples the accuracy of recovered equations compared to using raw data, while incurring negligible computational cost and requiring no additional training. The method enables language models to read field data as a theorist would, facilitating automated field‑theory construction that can evolve alongside experimental data.
By Fan Yang, Matt Thomson
arXiv:2603. 11479v3 Announce Type: replace-cross Abstract: Time Series Event Detection (TSED) aims to localize semantically meaningful events in time series data, with critical applications in high-stakes domains.
By Sky Chenwei Wan, Yifei Y. Wang, Tianjun Hou, Xiqing Chang, Aymeric Jan
The paper introduces SciGram, a large-scale dataset of 194K scientific diagrams paired with 1.4M visual instructions generated through a terminology‑grounded pipeline that extracts domain concepts, synthesizes facts, and retrieves relevant diagrams. Models fine‑tuned on SciGram show significant gains on diagram‑centric benchmarks such as TQA, ScienceQA, and AI2D, and when combined with existing models like LLaVA OneVision, set new state‑of‑the‑art performance. The authors release both the dataset and trained models to support further research in scientific diagram understanding.
By Raul Ortega, Jos\'e Manuel G\'omez-P\'erez
arXiv:2607. 03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation.
By Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang
SciDocBench is a workflow-centered benchmark for scientific document understanding that includes 124 expert-authored questions across seven capability groups and 19 subtasks in five scientific domains. Each question is evaluated under four conditions—English or Chinese, all-images-first or interleaved document representations—resulting in 496 evaluation instances. The benchmark is paired with SciDocIR, a typed evidence-graph representation, and SciDocDataset, a collection of 15K fine-tuning and 8K reinforcement-learning samples, forming an evaluation-to-training framework for scientific-document assistants.
By Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin
arXiv:2605. 29475v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery.
By Hongran An, Zonglin Yang