arXiv AI

ORCA: Evaluating LLMs on Data Science Code Translation

ORCA is a new benchmark for evaluating large language models on Data Science Code Translation (DSCT), comprising two settings: ORCA-MAIN with 1,600 grounding-level tasks across data querying, manipulation, and deep learning, and ORCA-PROJECT with 200 full-project translation tasks across seven data‑science task types. Each task includes reference translations and test cases to verify functional equivalence, and a multi‑stage quality verification process ensures task correctness. Experiments show that even state‑of‑the‑art LLMs perform poorly on DSCT, with Claude‑Opus‑4.6 achieving only 56.92% success on ORCA‑MAIN and 33.67% on ORCA‑PROJECT, while an intent‑augmented approach improves success rates by 4.80% and 5.33% respectively.

arXiv AI
Jun 6

Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation

arXiv:2512. 03086v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown remarkable capabilities in code translation, yet their performance deteriorates in low-resource programming domains such as Fortran and emerging frameworks like CUDA, where high-quality parallel data are scarce.

By Le Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao
arXiv AI
Sep 15

CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback

CHAI for LLMs is a framework that improves large language models’ performance on code‑mixed translation tasks by using LLMs as annotators to create preference data, applying reinforcement learning from AI feedback, incorporating LLM‑generated domain knowledge for iterative refinement, and evaluating on real‑world datasets. The approach yields a 68.45% average win rate over state‑of‑the‑art open‑source models in human‑adjudicated tests. It demonstrates a scalable method to enhance code‑mixed language understanding in open‑source LLMs.

By Wenbo Zhang, Aditya Majumdar, Asif Ekbal, Amulya Yadav
arXiv Machine Learning
Aug 31

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

The paper introduces CodeRQ-Bench, the first benchmark for assessing large language model reasoning quality across coding tasks such as generation, summarization, and classification. It analyzes over a thousand mismatches from existing evaluators, identifies recurring limitations, and derives design insights that lead to a new two‑stage evaluator, VERA. Experiments show VERA outperforms strong baselines, improving AUCROC by up to 0.26 and AUPRC by up to 0.21 on four datasets.

By Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed
arXiv AI
Sep 25

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

SciWalker is a framework that automatically synthesizes scientific coding problems by sampling operator chains from scientific library interfaces and using execution feedback to refine generated problem statements, solutions, and tests. It produces 8,178 high‑quality problems across five scientific domains and 32 subdomains, and training a large language model with these problems improves its scientific coding accuracy by nearly 10 percentage points. The approach combines structured workflow composition with verification and quality review to enable scalable, scientifically grounded task generation.

By Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye, Yu Cheng, Jun Zhang
arXiv AI
Aug 26

Evaluating Language Models on Cross-Language Code Functional Equivalence

The paper introduces PolyHuman, a dataset of human-written programs in C++, Java, and Python, to test whether large language models can judge functional equivalence across languages. Using this dataset, the authors evaluate several open-weight and proprietary LLMs, finding that models struggle more with harder problems, show language-specific biases, and rely partly on superficial similarity cues. They also observe run‑to‑run instability in GPT‑o4‑mini, concluding that current LLMs do not reliably capture functional equivalence within or across programming languages.

By Hui Sun, Anderson Uch\^oa, Rohit Gheyi, Wesley K. G. Assun\c{c}\~ao