Semantic Navigation for Issue Localization in Code Repository
arXiv:2609. 31176v1 Announce Type: new Abstract: Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue.
arXiv:2508. 12232v5 Announce Type: replace-cross Abstract: Issue-to-commit link recovery in software repositories is fundamental to software traceability and project management, yet it remains a challenging task.
arXiv:2609. 31176v1 Announce Type: new Abstract: Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue.
arXiv:2507. 16395v3 Announce Type: replace Abstract: Atomic commits, which address a single development concern, are a best practice in software development.
StateTape introduces a new framework for long‑horizon coding agents that rewrites the agent’s context as the code repository changes, rather than letting the context grow with every observation. It models the repository as a symbol‑level code graph, using a tape to mark symbols altered by each write and a manager model to resolve stale records. The authors provide theoretical analysis, a new benchmark called TraceBench, and empirical results showing higher resolve rates across six agents and three edit‑heavy benchmarks with minimal computational overhead.
arXiv:2606. 11976v1 Announce Type: cross Abstract: Software engineering tools increasingly rely on LLM based agents to localize files to change to resolve a software issue.
arXiv:2607. 13111v1 Announce Type: cross Abstract: Distinguishing semantic-preserving commits from changing ones remains an open challenge in software repository mining.
arXiv:2605. 17965v2 Announce Type: replace-cross Abstract: Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR).
arXiv:2605.14563v3 Announce Type: replace-cross Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and...
arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.
arXiv:2609.01601v1 Announce Type: cross Abstract: The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repo...
Software engineering tools increasingly rely on LLM based agents to localize files to change to resolve a software issue. Most AI agents explore repositories linearly, that is, visiting one directory or file per step.
The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.
SpecMine is a large-scale corpus that documents Spec-Driven Development (SDD) artifacts in public GitHub repositories. It includes a broad census of 470,795 spec files from 73,030 repositories linked to 17 tools, a focused census of 98,574 Kiro layout files from 12,910 repositories, and a sweep of 5,992 pull requests across 581 repositories that modify specs. The dataset provides enriched metadata, full commit histories, parsed document structures, and over 2.4 million typed references connecting specs to code, sibling documents, PRs, branches, and issues.