arXiv Machine Learning

Fixed-Set Robustness in Programming by Example: Example Corruption and Semantic Partition Recovery

arXiv:2607. 01280v1 Announce Type: new Abstract: Programming-by-example systems infer programs from a small set of input-output examples.

Hugging Face Trending Papers
Jun 2

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate computation to code. Yet models are still often used in settings where they must reason directly from natural language, and trustworthy models should solve small-number arithmetic word problems without external tools.

arXiv AI
Aug 24

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

The paper introduces Trustworthy RAG, an evaluation agent designed to detect misinformation and knowledge poisoning in Retrieval-Augmented Generation systems. It combines natural language inference verification, a five-signal poison detector, and a weighted Trust Index to assess the reliability of retrieved content. Experiments on multiple LLMs show high accuracy and precision, with the agent effectively blocking unsafe advice in a secure-coding assistant scenario.

By Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
arXiv AI
Aug 20

A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

The paper investigates how AI code agents perform when the surrounding code is rewritten in a semantically equivalent way. Using a random variant sampler that applies control‑flow rewrites, dead‑code injection, and identifier renaming, the authors evaluate two agent scaffolds—mini‑SWE agent and OpenCode—backed by four frontier models across SWE‑bench datasets. Results show modest drops in resolve‑rate (up to 6.7 percentage points) with significant degradations in 6 of 16 configurations, and reveal that robustness varies across models and scaffolds, forming a jagged frontier.

By Hasan Najib Mahmud (Colorado State University), Shreya Gupta (Microsoft), Isha Chaudhary (University of Illinois Urbana-Champaign), Nathaniel Enis (Colorado State University), Ravi Mangal (Colorado State University), Gagandeep Singh (University of Illinois Urbana-Champaign), Corina Pasareanu (Carnegie Mellon University)
arXiv Computation and Language
Sep 25

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how large language models (LLMs) handle bug fixing versus problem solving in competitive programming. Using a dataset of ~3,000 Codeforces submissions and their human fixes, the authors compare LLM-generated patches to human patches and assess whether LLMs prefer to modify buggy code or generate new solutions. Results show that LLMs often alter more lines than necessary and sometimes produce entirely new solutions, performing better when allowed to generate solutions from scratch rather than patching existing code.

By Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu
arXiv Machine Learning
Sep 10

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

The paper evaluates how robust large language models are at generating SystemVerilog Assertions (SVA) when the underlying RTL code undergoes semantics‑preserving transformations such as operand reordering, identifier renaming, and redundant parenthesization. Using a curated dataset and two open‑source models (Qwen2.5‑Coder‑7B and DeepSeek‑Coder‑V2‑Lite), the authors find that 9.7%–27.0% of behaviors that were correct on the original RTL become incorrect after transformation, revealing significant instability that aggregate accuracy metrics can hide.

By FNU Aditi
arXiv AI
Sep 12

Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code

The paper introduces the Static‑Pass Dynamic‑Fail (SPDF) phenomenon, showing that static analysis can miss vulnerabilities that are exploitable at runtime. Using a three‑stage pipeline—static scanning, LLM‑driven CWE reasoning, and autonomous exploit verification—it evaluated 1,355 Python samples and found that 14.53% of samples that passed static checks were actually exploitable. The study highlights that static‑analysis success and runtime security are distinct assurance layers, especially for AI‑generated and security‑sensitive code.

By Jessica Pourleyli, Maitreyee Das Urmi, Glaucia Melo
arXiv AI
Sep 21

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...

By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras