arXiv AI By Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini

Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

Read the original on arXiv AI →

The paper introduces FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit continuous integration workflow. FlowAgent uses a ReAct-style generate-and-validate loop with strict latency and quality filters, and was evaluated on 195 real-world failures with a 67.18% accuracy rate. After deployment, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554, and received positive feedback from interviews.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study

The paper presents a systematic analysis of five state‑of‑the‑art automated program repair agents, tracing their decision‑making across 500 real‑world repair tasks. It finds that while the agents perform well on simple fixes, they struggle with logic‑intensive bugs, often producing verbose, overfitted patches that pass tests without addressing root causes. Key bottlenecks identified include poor test generation, limited regression test selection, and reliance on primitive tooling without access to debuggers or advanced program analysis tools.

By Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
arXiv AI
1d ago

AegisFlow: A Multi-Agent Agentic AI Framework for Autonomous Remediation and Self-Healing in Fragile Data Ecosystems

AegisFlow is a multi‑agent AI framework that automates the detection and remediation of failures in data pipelines, using a Watchdog agent for telemetry and a Repair agent that generates, tests, and deploys code patches via large language models. It employs a non‑intrusive Parallel Shadow Patching approach based on the MAPE‑K loop to validate patches in digital twin environments. Experiments across five failure scenarios show a 98.1% reduction in mean time to repair—from 170 minutes to 3.2 minutes—and a 92% patch success rate, freeing up 98% of data engineering on‑call time for innovation.

By Muhammad Bilal Awan, Zubair Hussain, Abdul Shahid