arXiv AI

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Dude is a Dual-Detection Multi-Agent System designed to detect discrepancies between research papers and their accompanying code. It addresses limitations of single-agent LLM approaches by aligning granularity between paper and code through negotiation and a two-stage salience-filtering mechanism, reducing false positives. Experiments on real-world datasets show that Dude improves recall and precision by up to 22.8% and boosts the F1 score by up to 18.7% over baseline methods.

Hugging Face Trending Papers
Jun 23

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly understood. We introduce a multi-layered detection framework that integrates configuration-file scanning, commit-message analysis, author-identity matching, and bot-signature lookup across World of Code (180M+ Git repositories), classifying agent traces into four behavioral types.

arXiv AI
6d ago

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.

By Md Shohel Arman, Igor Molybog
arXiv AI
Jul 3

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

arXiv:2607. 01916v1 Announce Type: new Abstract: Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs.

By Chiwang Luk, Matin Mohammad Najafi, Zhifeng Jia, Wei Yang, Xiuchang Li, Jinwei Zhu, Yang Ren, Lei Chen, Gao Cong