SciWalker is a framework that automatically synthesizes scientific coding problems by sampling operator chains from scientific library interfaces and using execution feedback to refine generated problem statements, solutions, and tests. It produces 8,178 high‑quality problems across five scientific domains and 32 subdomains, and training a large language model with these problems improves its scientific coding accuracy by nearly 10 percentage points. The approach combines structured workflow composition with verification and quality review to enable scalable, scientifically grounded task generation.
By Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye, Yu Cheng, Jun Zhang
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
By Ankit Gupta, Aditya Prasad, Rameswar Panda
ORCA is a new benchmark for evaluating large language models on Data Science Code Translation (DSCT), comprising two settings: ORCA-MAIN with 1,600 grounding-level tasks across data querying, manipulation, and deep learning, and ORCA-PROJECT with 200 full-project translation tasks across seven data‑science task types. Each task includes reference translations and test cases to verify functional equivalence, and a multi‑stage quality verification process ensures task correctness. Experiments show that even state‑of‑the‑art LLMs perform poorly on DSCT, with Claude‑Opus‑4.6 achieving only 56.92% success on ORCA‑MAIN and 33.67% on ORCA‑PROJECT, while an intent‑augmented approach improves success rates by 4.80% and 5.33% respectively.
By Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng
arXiv:2603. 14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones.
By Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, Li Zhang
arXiv:2511. 06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce runtime while preserving program correctness.
By Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh, Kevin Swersky, Ofir Press, Enhui Li, Vijay Janapa Reddi, Parthasarathy Ranganathan
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
By Jie Li, Wenzhao Wu, Junqi Hu, Qinrui Zheng, Bowen Wu, Juepeng Zheng, Yutong Lu, Haohuan Fu
arXiv:2607. 05443v1 Announce Type: cross Abstract: Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging.
By Nishan Pantha, Pranath Reddy Kumbam, Sajil Awale, Pushwitha Krishnappa, Muthukumaran Ramasubramanian, Nidhi Jha, Emily Foshee, Ankur Kumar, Rachel Slank, Ashkbiz Danehkar, Rahul Ramachandran
arXiv:2608. 04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code.
By Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen
arXiv:2603. 16011v3 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to optimize entire codebases under realistic constraints.
By Atharva Sehgal, James Hou, Akanksha Sarkar, Ishaan Mantripragada, Swarat Chaudhuri, Jennifer J. Sun, Yisong Yue
arXiv:2609.12475v1 Announce Type: new
Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluati...
By Zhongzhan Huang, Junxin Li, Guoming Ling, Yupei Lin, Shanshan Zhong, Hefeng Wu
SciMIF is a new benchmark that evaluates how well multimodal large language models (MLLMs) can follow complex scientific instructions. It is built on an analysis of 22 tasks across five scientific fields and introduces a taxonomy of 10 constraint groups that capture both general and discipline‑specific requirements. Experiments show large performance gaps between fields—chemistry is hardest—and that larger models do not necessarily improve constraint adherence, especially for fine‑grained, knowledge‑heavy instructions.
By Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai
The survey reviews how Large Language Models (LLMs) are being used in High‑Performance Computing (HPC) programming, covering code generation, parallelization, frameworks, evaluation, and broader challenges. It finds that general‑purpose LLMs perform adequately on serial and OpenMP‑style tasks but struggle with distributed MPI workloads, while domain‑specialized models achieve higher accuracy yet are limited in scope and evaluation. The authors argue that LLMs will not replace HPC experts soon but can act as powerful collaborators, provided richer datasets, integration with performance tools, rigorous evaluation, and governance are developed.
By Strahinja Ljaljevic, Josep Jorba, Sergio Iserte