arXiv AI

Loreley: Repository-Scale Program Evolution with Quality-Diversity Search

arXiv:2608. 19703v1 Announce Type: cross Abstract: Sequential agent search accumulates changes from its current champion but discards alternative branches; independent proposals preserve breadth but restart from the root.

arXiv AI
Sep 17

Agora: Git as Shared Memory for Collective AutoResearch

Agora is a system that uses Git as a shared memory for autonomous research agents, recording each claim as an immutable commit in an append‑only directed acyclic graph. In a 12‑day run, 13 language‑model workers independently explored a weight‑transfer problem, producing 1,703 contributions that improved a 119.6M‑parameter model’s performance from 3.39 to 1.899 bits per byte. The system’s design includes a diversity‑aware selection rule and an index that tracks the frontier, neglected branches, and verification status of each claim.

By Yifan Zhang, Yunheng Zou, Shaokun Zhang, Jian Hu, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong
arXiv AI
Sep 2

REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.

By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
arXiv Machine Learning
Sep 23

Impact Is Not Invalidation: Ask About the Claim, Not the Diff

The paper investigates how machine‑learning models can determine whether a claim (a test assertion) remains valid after a code change. It compares two questioning strategies: asking whether a diff preserves behavior versus asking whether a specific claim still holds. The authors find that the latter approach yields far higher precision (up to 0.974) across models of varying cost, while the former performs poorly (precision 0.291–0.329). They also benchmark against a regression‑test selector, showing that even near‑complete knowledge of a change’s reach does not reliably identify falsified claims. The study is grounded in 10,369 mined claims with 184 execution‑verified flips from 23 Python libraries.

By Atul Anand
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv AI
Sep 4

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

The paper reports on a series of experiments examining how different forms of directives—such as record pointers, criteria, or combinations—affect an agent’s choice of archived source records when it inherits six one-line memories. Across twelve registered studies involving 14,760 attempts on a single instrument lineage, the authors measured the impact of various directive formats on six direct-provider models, nine OpenRouter-served models, and several Claude and Opus 5 models, noting differences in performance metrics and replication outcomes. The results are purely descriptive, detailing the effects of exact edits on fixed panels with registered intervals and no claim of underlying mechanisms.

By Kazuki Nakayashiki
arXiv AI
Sep 1

Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

The study investigates how inference‑time interventions and weight consolidation affect open‑ended generation in an online bin‑packing task. By iteratively generating, verifying, selecting, and consolidating with LoRA, the model’s outputs shift toward higher value, reducing excess by 1.7 points and outperforming random consolidation by 3.1 points. Across three independent runs, the mean performance remained consistent, and the best candidates converged to the classic heuristic’s level without exceeding it, while consolidation also lowered the proportion of better‑than‑classic candidates but increased their absolute number.

By Roberto I. Ono Filho
arXiv AI
Aug 25

Repo2Skill-Evo: Repository Skills Go Stale in Silence

arXiv:2608.21964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural kno...

By Chenyuan Duan, Ge Shi, Zineng Mao, Ge Zhang, Hao Liang, Yinzhu Piao, Yuchen Wu, Zhixin Yao, Kaiyu Huang, Wenhao Huang, Linzhuang Sun, Shen Yan, Wentao Zhang