arXiv AI

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

SciWalker is a framework that automatically synthesizes scientific coding problems by sampling operator chains from scientific library interfaces and using execution feedback to refine generated problem statements, solutions, and tests. It produces 8,178 high‑quality problems across five scientific domains and 32 subdomains, and training a large language model with these problems improves its scientific coding accuracy by nearly 10 percentage points. The approach combines structured workflow composition with verification and quality review to enable scalable, scientifically grounded task generation.

arXiv AI
Jul 28

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

arXiv:2509. 03059v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where ground-truth correctness can be automatically evaluated.

By Xingyue Huang, Rishabh, Gregor Franke, Ziyi Yang, Jiamu Bai, Weijie Bai, Jinhe Bi, Zifeng Ding, Yiqun Duan, Chengyu Fan, Wendong Fan, Xin Gao, Ruohao Guo, Yuan He, Zhuangzhuang He, Xianglong Hu, Neil Johnson, Bowen Li, Fangru Lin, Siyu Lin, Tong Liu, Yunpu Ma, Hao Shen, Hao Sun, Beibei Wang, Fangyijie Wang, Hao Wang, Haoran Wang, Yang Wang, Yifeng Wang, Zhaowei Wang, Ziyang Wang, Yifan Wu, Zikai Xiao, Chengxing Xie, Fan Yang, Junxiao Yang, Qianshuo Ye, Ziyu Ye, Guangtao Zeng, Yuwen Ebony Zhang, Zeyu Zhang, Zihao Zhu, Bernard Ghanem, Philip Torr, Guohao Li
arXiv AI
Jun 4

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

arXiv:2606. 04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as biology, chemistry, and physics remains largely unexplored.

By Xiangyu Zhao, Hengyuan Zhao, Yiheng Wang, Wanghan Xu, Yuhao Zhou, Qinglong Cao, Zhiwang Zhou, Lei Bai, Wenlong Zhang, Xiao-Ming Wu
Hugging Face Trending Papers
Jul 29

SciDataSailor: Deep Scientific Data Exploring

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.

arXiv Machine Learning
Aug 27

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.

By Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai
arXiv AI
6d ago

SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning

SLMFix is a code‑generation pipeline that uses a small language model fine‑tuned with reinforcement learning to correct syntactic errors in programs produced by large language models for domain‑specific languages. The approach relies on interpreter feedback to guide the error‑fixing process. Experiments show that SLMFix improves validator pass rates by 40% on low‑resource programming languages and removes over 50% of syntactic errors on high‑resource DSLs, outperforming supervised fine‑tuning even for 7B models.

By David Jiahao Fu, Aryan Gupta, Aaron Councilman, Yu-Xiong Wang, Vikram Adve
arXiv AI
6d ago

ORCA: Evaluating LLMs on Data Science Code Translation

ORCA is a new benchmark for evaluating large language models on Data Science Code Translation (DSCT), comprising two settings: ORCA-MAIN with 1,600 grounding-level tasks across data querying, manipulation, and deep learning, and ORCA-PROJECT with 200 full-project translation tasks across seven data‑science task types. Each task includes reference translations and test cases to verify functional equivalence, and a multi‑stage quality verification process ensures task correctness. Experiments show that even state‑of‑the‑art LLMs perform poorly on DSCT, with Claude‑Opus‑4.6 achieving only 56.92% success on ORCA‑MAIN and 33.67% on ORCA‑PROJECT, while an intent‑augmented approach improves success rates by 4.80% and 5.33% respectively.

By Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng
arXiv Computation and Language
Sep 17

Code Consistency Preference Optimization Verification for Language Model Alignment

The paper introduces Code Consistency Preference Optimization Verification (CCPO), a method that generates computationally sound solutions with dependency graphs to improve execution-consistent preference optimization for language models. By building a scientific reasoning dataset and extracting reasoning steps, prerequisites, and derivability relationships, the authors compute execution consistency scores that are used to fine‑tune models such as Llama‑3‑8B and DeepSeekMath‑7B, achieving significant performance gains on MATH (+17.0%) and GSM8K (+15.1%). The extended Scientific Feasibility Control framework further boosts accuracy on PhyX physics reasoning to 50.1%, surpassing existing models while maintaining high scientific validity and reducing law violations.

By Yunlong Tan, Mingqiao Mo, Hao Zhang