arXiv AI

BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs

arXiv:2608. 13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve.

arXiv AI
Sep 25

The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information

The paper investigates how large language model agents decide whether to persist, stop, or escalate when faced with impossible software‑repair tasks that also involve conflicting test requirements. Using ImpossibleBench tasks and models such as GPT‑5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash, the study varies peer precedent, forged authority claims, instruction wording, and tool friction to observe differing adjudication policies. The authors propose a conflict adjudication framework that maps information to interpretation to action, arguing it better captures agent alignment under competing pressures.

By Ivy Zhang
arXiv AI
Sep 12

Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

The paper argues that deploying generative AI agents requires more than isolated task success; they must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows. The authors introduce two complementary evaluation aspects—operational resilience and considerate participation—to assess how agents recover from blocked work, communicate limits, and adapt to affected people and role boundaries. Using 120 simulated healthcare trajectories across two AI models and twelve stakeholder-derived tasks under varying challenge levels, the study finds that agents shift toward greater human dependence and increased workload as challenge accumulates, while also broadening from task-focused adaptation to task reframing and wider coordination.

By Yuanchen Bai, Zijian Ding, Angelique Taylor
arXiv AI
Sep 10

Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction

The paper introduces Graph‑Agentic Retrieval‑Augmented Generation (RAG), a system that blends structured evidence with adaptive agents capable of planning retrieval, navigating relations, verifying claims, delegating tasks, and employing tools. It highlights how defects in graph construction can propagate through retrieval and control decisions, potentially leading to significant outcomes. To address these risks, the authors propose an assurance‑by‑construction framework with five interface contracts—evidence, retrieval, reasoning, capability & delegation, and outcome—that make provenance, validity, authorization, uncertainty, and recoverability explicit, and outline an evaluation agenda for social‑good applications.

By Vijay Bommireddy, Raviteja Bommireddy
arXiv AI
Aug 20

One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI

The paper investigates how multiple pre‑action controls—authority, resource, and evidence gates—interact in agentic AI systems. It formalizes remediation‑induced control coupling, showing that remediation can invalidate earlier judgments and that the order of remediation matters. The authors propose a remediate‑and‑regate protocol to restore soundness, analyze non‑commuting remediation operators, and demonstrate the approach on a deterministic open‑data artifact with three published engines.

By Gaston Besanson
arXiv AI
Sep 21

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...

By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)
arXiv AI
Sep 10

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

The paper introduces Procedural Graphs, a framework that structures procedural knowledge for large language model agents as (procedure, relation, procedure) triplets, analogous to knowledge graphs for factual data. At each decision point, a guidance model uses the local subgraph to bias the agent’s next action, while an LLM refiner self‑evolves the graph by comparing failed and successful trajectories, editing its topology to improve performance. Experiments across various datasets, tasks, and LLMs show that Procedural Graphs consistently outperform memory‑based baselines, and the self‑evolution mechanism further enhances results without manual engineering.

By Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan \"{O}. Ar{\i}k
arXiv AI
Aug 28

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

The paper reports a failure study of a production agentic software‑delivery platform, analyzing 147 incidents across 81 runs. It shows that the standard reliability primitives—retry, timeout, and error‑rate circuit breaking—fail in practice, leading to costly loops, false trips, and blocked work. The authors identify two cross‑cutting causes—identity adequacy and evidence adequacy—and propose seven new reliability primitives that enforce reliability at the delegation level.

By Mazhar Shaikh, Anurag Rajkumar Bombarde, Harshal Pathak
arXiv Machine Learning
1d ago

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.

By Elias Malomgr\'e, Pieter Simoens