arXiv:2607. 19104v1 Announce Type: cross Abstract: Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question.
By Weifeng Sun, Ye Fan, Yuchen Chen, Gou Tan, Jieke Shi, Yuan Yidi, Swee Liang Wong, Jonathan Pan, David Lo
arXiv:2507. 11687v5 Announce Type: replace-cross Abstract: Large language models excel at code generation but struggle with code linting, particularly in generalizing to unseen or evolving best practices beyond those observed during training.
By Atharva Naik, Lawanya Baghel, Dhakshin Govindarajan, Darsh Agrawal, Yiqing Xie, Daniel Fried, Carolyn Rose
arXiv:2608. 04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code.
By Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen
arXiv:2509. 20491v3 Announce Type: replace-cross Abstract: Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code.
By Brahim Mahmoudi, Naouel Moha, Quentin Sti\'evenart, Florent Avellaneda
arXiv:2607. 04537v1 Announce Type: cross Abstract: Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benchmark that evaluates them assumes the instructions they act on are correct.
By Raj Jaiswal, Anany Singh Divy, Savar Bhasin, Adi Bajpai, Tanuja Ganu, Rajiv Ratn Shah
Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benchmark that evaluates them assumes the instructions they act on are correct. We study what happens when that assumption breaks.
arXiv:2606. 29088v1 Announce Type: cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models.
By Bal\'azs Szalontai, \'Abel Szauter, Bal\'azs M\'arton, P\'eter Verebics, Bal\'azs Pint\'er, Tibor Gregorics
ToolGate is an executable acceptance pipeline designed to streamline the creation of scientific benchmarks that rely on specialist software. It evaluates each model-generated item through three gates: (1) an executable solution script must reproduce the proposed answer, (2) a randomized no‑tool screen rejects items solvable without the software, and (3) a tool‑using agent must solve the item within a time limit. In a FEniCSx instantiation, 500 generation attempts produced 128 unique, verified benchmark items after successive filtering.
By Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.
By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang
arXiv:2606. 23877v1 Announce Type: cross Abstract: Jupyter Notebooks are an increasingly popular coding environment used across many domains, especially in Python-based data science and scientific computing.
By Lukas Ottenhof, Thibaud Lutellier
arXiv:2607. 28587v2 Announce Type: replace-cross Abstract: SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability.
By Manyi Wang, Junjielong Xu, Pinjia He
arXiv:2608. 16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can select and tune known components, but cannot produce structure outside that space.
By Sofoklis Kitharidis, Cor J. Veenman, Jan N. van Rijn, Thomas B\"ack, Niki van Stein