arXiv Computation and Language

Taramandal-GPT: Enhancing Astrodynamics Problem-Solving with Knowledge Retrieval and Structured Thinking

arXiv AI
Sep 17

Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

The paper investigates whether domain-specific fine‑tuning benefits open‑ended scientific reasoning in astronomy. Using a curated 300‑question QA benchmark from 2017–2026 Olympiad‑style materials, the authors compare open‑weight, API‑served general‑purpose, multimodal, and astronomy‑specialized language models. Results show that strong general‑purpose models set the highest correctness baseline, but variations in metric agreement, judge sensitivity, benchmark composition, and modality suggest that domain specialization is task‑ and deployment‑dependent and that domain‑specific evaluation is crucial for scientific workflows.

By Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang
arXiv Computation and Language
Aug 31

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

NL2AGBench is a benchmark that evaluates how well large language models can translate English geometry problems into the formal language required by AlphaGeometry’s theorem‑proving engine. The study tests ten state‑of‑the‑art LLMs, comparing executable translation accuracy, syntactic correctness, and error types, and finds a large gap between closed‑source and open‑source models. The authors also propose an error taxonomy and test mitigation strategies such as few‑shot prompting, fine‑tuning, and human‑guided hinting, which improve performance across model families.

By Samuel Xiao, Judy Song, Rory Hu, Ziliang Zong
arXiv Machine Learning
Sep 3

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

The paper presents a SciBERT-based method for automatically classifying scientific papers into four telescope-related categories—science, instrumentation, mention, and not telescope—within strict 512-token limits. Despite truncation challenges, the approach achieved a macro F1 score of 0.89, topping the WASP-2025 leaderboard. The authors analyze truncation effects, compare chunking and long-context models, and offer insights into efficient scientific text curation.

By Madhusudhana Naidu
Hugging Face Trending Papers
Jul 29

SciDataSailor: Deep Scientific Data Exploring

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.

arXiv AI
Sep 18

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.

By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan
arXiv Computation and Language
Sep 1

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

arXiv:2606.15872v2 Announce Type: replace Abstract: Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of...

By Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin