AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.24246v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable progress in natural language understanding, yet their effectiveness in specialized fields like astro...
The paper investigates whether domain-specific fine‑tuning benefits open‑ended scientific reasoning in astronomy. Using a curated 300‑question QA benchmark from 2017–2026 Olympiad‑style materials, the authors compare open‑weight, API‑served general‑purpose, multimodal, and astronomy‑specialized language models. Results show that strong general‑purpose models set the highest correctness baseline, but variations in metric agreement, judge sensitivity, benchmark composition, and modality suggest that domain specialization is task‑ and deployment‑dependent and that domain‑specific evaluation is crucial for scientific workflows.
The paper presents a SciBERT-based method for automatically classifying scientific papers into four telescope-related categories—science, instrumentation, mention, and not telescope—within strict 512-token limits. Despite truncation challenges, the approach achieved a macro F1 score of 0.89, topping the WASP-2025 leaderboard. The authors analyze truncation effects, compare chunking and long-context models, and offer insights into efficient scientific text curation.
An agentic framework called GW‑Eyes, powered by large language models, is introduced to autonomously associate gravitational‑wave (GW) signals with candidate electromagnetic (EM) counterparts. It integrates domain‑specific tools for tasks such as catalog management, skymap visualization, and rapid verification, while enabling natural‑language interaction to assist human experts. The framework leverages LLMs’ decision‑making and traceable reasoning to address the growing data‑analysis challenges of next‑generation GW and EM detectors.
SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.
arXiv:2609.13648v1 Announce Type: new Abstract: Solar energy decision support is fragmented across dashboards that provide data without explanation, research papers are slow to parse, and general-pur...