arXiv:2607. 19104v1 Announce Type: cross Abstract: Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question.
By Weifeng Sun, Ye Fan, Yuchen Chen, Gou Tan, Jieke Shi, Yuan Yidi, Swee Liang Wong, Jonathan Pan, David Lo
arXiv:2608. 19799v1 Announce Type: new Abstract: Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions.
By Zhipeng Xu, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu
arXiv:2608. 05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review.
By Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
arXiv:2609.00654v1 Announce Type: new
Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scien...
By Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan
SWE‑bench Science is a benchmark for evaluating coding agents on scientific software engineering tasks, comprising 119 tasks from 98 GitHub repositories across 20 scientific domains. The tasks are grouped into Issue‑driven, Expert‑exploratory, and Engineering‑integration paradigms, and even the best agent, Claude Code with Opus‑5 (max), achieves a pass@1 below 50%. The study identifies four common failure mechanisms—lack of scientific knowledge, misguided exploration, incomplete repair coverage, and poor generalization—and shows that providing scientific guidance can both help and hinder repair depending on its alignment with the task.
arXiv:2606. 29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof.
By Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman
The paper introduces Rules to Tools (R2T), a system that provides executable checks for scientific coding agents to verify compliance with public scientific requirements. In experiments across multiple task cohorts, agents using R2T’s prepared checks achieved high repair success rates—26/30 with text and 29/30 with checks—while also demonstrating varying task preferences and cost trade‑offs. The study quantifies how tool‑enabled checks influence repair outcomes and agent‑side resource usage in scientific computing contexts.
By Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke
The paper introduces Lit2Test, a benchmark that evaluates language models’ research idea proposals by requiring each idea to include a falsifiable outcome, thereby making quality decidable. Built from 200 real-paper neighborhoods, the benchmark gathers proposals from four frontier models and compares them via 1,200 blind pairwise judgments, with reliability checks and human calibration. The results show a consistent ranking of the models, driven by test and metric quality rather than fluency, and the authors release the benchmark and related artifacts for public use.
By Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Hongyao Zuo (Tianjin University), Ziwen Gong (Hainan University), Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Tong Yang (Peking University), Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Xiaohui Li (Huawei Technologies), Haoli Bai (Huawei Technologies)
arXiv:2606. 12864v1 Announce Type: cross Abstract: Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored.
By Tingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu, Kaifeng Lyu
arXiv:2609.21190v1 Announce Type: cross
Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...
By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
arXiv:2609.12708v2 Announce Type: replace-cross
Abstract: AI coding assistants are becoming co-authors of production software, yet their evaluation centers on functional correctness, leaving open whe...
By Cristina Improta, Pietro Liguori, Domenico Cotroneo
arXiv:2605. 26548v2 Announce Type: replace-cross Abstract: Finding a real vulnerability in complicated systems is a challenging, long-horizon task that demands reasoning across an entire codebase to produce a working proof-of-concept (PoC).
By Hwiwon Lee, Jiawei Liu, Dongjun Kim, Wubing Xia, Ziqi Zhang, Chunqiu Steven Xia, Lingming Zhang