arXiv:2606. 10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogged, diagnoses cannot be replayed, and promote-or-discard decisions land in a side database rather than the agent's own history.
By Yohei Nakajima
The paper introduces Revision‑Aware Independent Agent Graphs (RIAG) to address dynamic task routing, where an event stream continually revises task bindings and a system must select the correct document version at query time. By repurposing six benchmarks into over 31,000 dynamic episodes, the authors demonstrate that RIAG balances recomputation and reuse, achieving 54.24 % joint routing‑and‑answer accuracy with only 0.62 calls per query—substantially better than the strongest baseline. The study highlights the trade‑off between stale conclusions and wasted work in dynamic reasoning settings.
By Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang
The paper introduces “Revise”, a runtime system that performs validity-guided, fine-grained recovery for online revisions in structured agent workflows. When a revision arrives, Revise intersects the change with recorded data and control dependencies, propagates the impact through the partially executed DAG, stops invalid work, preserves unaffected progress, and recomputes only the affected region. Experiments on real coding‑agent traces and LangGraph/LLMCompiler applications show that Revise matches a latest‑version oracle, reduces model calls by up to 56%, and improves service‑level objective goodput under load.
By Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
arXiv:2608. 14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures.
By Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li
arXiv:2607. 16345v1 Announce Type: cross Abstract: Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task.
By Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou
LEDGER is a tracing and review system for large language model agents that constructs layered trace graphs from observed sessions. It groups raw trace records into Evidence Nodes and Workflow Nodes, anchors artifacts as evidence, and adds typed semantic edges linking claims to supporting actions, artifacts, and checks. The resulting traces reveal workflow decisions, artifact lineage, repair steps, validation coverage, and claim‑support paths for evidence‑centered audit.
By Daehong Kim, Haichao Miao, Shusen Liu
arXiv:2607. 09682v1 Announce Type: new Abstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare.
By Vimal Nakrani
arXiv:2602. 17990v2 Announce Type: replace Abstract: Multi-agent LLM systems that generate structured workflows from natural-language requests are now deployed in production across cloud automation, DevOps, and enterprise process orchestration.
By Madhav Kanda, Sharad Agarwal, Rodrigo Fonseca, Alok Gautam Kumbhare, Pedro Las-Casas
EffectMatch is a runtime system that monitors and validates persistent changes made by large language model agents during software interactions. It collects all side‑effects within a controlled execution boundary and compares them against the application’s approved state, deciding whether to commit the changes and allow subsequent steps. In tests on 206 public business tasks, EffectMatch preserved correct executions and prevented all incorrect commits, with ablation studies showing the importance of each component.
By Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo, Hongzhi Wang
PinSieve is a production system that selectively serves vision‑language models (VLMs) for enterprise content‑quality triage, operating only on the grey‑zone cases that lightweight models cannot resolve. The deployed VLM Serving Agent filters 2.05× more non‑actionable items, improves review productivity by 25.7%, cuts operating costs by 16.2%, and delivers signals the same day instead of the next. A governed memory flywheel with selective feedback, audit sampling, and a bounded proposal‑verifier loop further reduces false‑negative rates from 17.73% to 13.29% over six months, while a reasoning review agent audits teacher‑generated rationales for keep/repair/drop decisions.
whyItMatters":"The system demonstrates how selective VLM serving and governed feedback loops can substantially improve efficiency, cost, and accuracy in enterprise AI content‑quality pipelines."
By Chuqing Gao, Yuanfang Song, Jonathan Zhang, Yifan Wu, Vishwakarma Singh, Qinglong Zeng, Andrey Gusev
The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.
By Elias Malomgr\'e, Pieter Simoens
arXiv:2608. 06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it.
By Shuyang Liu, Saman Dehghan, Ji Young Kim, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand