arXiv AI By Jinghang Shi, Yanxia Zhang, Ali Luo, Changhua Li, Xiao Kong

AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 17

Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

The paper investigates whether domain-specific fine‑tuning benefits open‑ended scientific reasoning in astronomy. Using a curated 300‑question QA benchmark from 2017–2026 Olympiad‑style materials, the authors compare open‑weight, API‑served general‑purpose, multimodal, and astronomy‑specialized language models. Results show that strong general‑purpose models set the highest correctness baseline, but variations in metric agreement, judge sensitivity, benchmark composition, and modality suggest that domain specialization is task‑ and deployment‑dependent and that domain‑specific evaluation is crucial for scientific workflows.

By Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang
arXiv Machine Learning
Sep 3

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

The paper presents a SciBERT-based method for automatically classifying scientific papers into four telescope-related categories—science, instrumentation, mention, and not telescope—within strict 512-token limits. Despite truncation challenges, the approach achieved a macro F1 score of 0.89, topping the WASP-2025 leaderboard. The authors analyze truncation effects, compare chunking and long-context models, and offer insights into efficient scientific text curation.

By Madhusudhana Naidu
arXiv AI
Sep 10

An agentic framework for gravitational-wave counterpart association in the multi-messenger era

An agentic framework called GW‑Eyes, powered by large language models, is introduced to autonomously associate gravitational‑wave (GW) signals with candidate electromagnetic (EM) counterparts. It integrates domain‑specific tools for tasks such as catalog management, skymap visualization, and rapid verification, while enabling natural‑language interaction to assist human experts. The framework leverages LLMs’ decision‑making and traceable reasoning to address the growing data‑analysis challenges of next‑generation GW and EM detectors.

By Yiming Dong, Yacheng Kang, Junjie Zhao, Xinyuan Zhu, Ziming Wang, Lijing Shao
arXiv Machine Learning
Aug 27

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.

By Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai