CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
arXiv:2603. 14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones.
arXiv:2603. 14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones.
arXiv:2512. 22827v2 Announce Type: replace-cross Abstract: Code often suffers from performance bugs.
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
arXiv:2605. 16046v2 Announce Type: replace-cross Abstract: Semantic code search has been widely adopted in both academia and industry.
arXiv:2505. 03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics.
arXiv:2607. 19104v1 Announce Type: cross Abstract: Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question.
arXiv:2506. 11066v3 Announce Type: replace-cross Abstract: Code retrieval is essential in modern software development, as it boosts code reuse and accelerates debugging.
The paper introduces PolyHuman, a dataset of human-written programs in C++, Java, and Python, to test whether large language models can judge functional equivalence across languages. Using this dataset, the authors evaluate several open-weight and proprietary LLMs, finding that models struggle more with harder problems, show language-specific biases, and rely partly on superficial similarity cues. They also observe run‑to‑run instability in GPT‑o4‑mini, concluding that current LLMs do not reliably capture functional equivalence within or across programming languages.
arXiv:2505.23126v5 Announce Type: replace Abstract: Although many benchmarks evaluate the reasoning abilities of Large Language Models (LLMs) within domains such as mathematics, coding, or data wrang...
The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.
arXiv:2606. 03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data.
arXiv:2507. 11687v5 Announce Type: replace-cross Abstract: Large language models excel at code generation but struggle with code linting, particularly in generalizing to unseen or evolving best practices beyond those observed during training.