Evaluating large language models trained on code
Related stories
Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc
arXiv:2607. 20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc, a constraint modeling language for combinatorial problems.
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
arXiv:2512. 20638v2 Announce Type: replace-cross Abstract: The evaluation of large language models relies heavily on standardized benchmarks.
FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis
arXiv:2607. 18260v1 Announce Type: new Abstract: We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis.
Matryoshka Language Model Suites
arXiv:2608. 09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently.
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
arXiv:2507. 11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset.
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming
arXiv:2606. 12864v1 Announce Type: cross Abstract: Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored.
M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs
arXiv:2606. 08840v1 Announce Type: new Abstract: Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance varies across programming languages, problem families, and failure modes.
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
arXiv:2605. 00754v4 Announce Type: replace-cross Abstract: Reward models (RMs) have become an indispensable fixture of the language model (LM) post-training playbook, enabling policy alignment and test-time scaling.
Improved Large Language Diffusion Models
arXiv:2606. 25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention.
Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions
arXiv:2606. 30790v1 Announce Type: cross Abstract: Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the dominant form of communication across multilingual communities.