arXiv:2608. 13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve.
By Sanjeev Manivannan
The paper introduces Graph‑Agentic Retrieval‑Augmented Generation (RAG), a system that blends structured evidence with adaptive agents capable of planning retrieval, navigating relations, verifying claims, delegating tasks, and employing tools. It highlights how defects in graph construction can propagate through retrieval and control decisions, potentially leading to significant outcomes. To address these risks, the authors propose an assurance‑by‑construction framework with five interface contracts—evidence, retrieval, reasoning, capability & delegation, and outcome—that make provenance, validity, authorization, uncertainty, and recoverability explicit, and outline an evaluation agenda for social‑good applications.
By Vijay Bommireddy, Raviteja Bommireddy
The paper investigates how large language model agents decide whether to persist, stop, or escalate when faced with impossible software‑repair tasks that also involve conflicting test requirements. Using ImpossibleBench tasks and models such as GPT‑5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash, the study varies peer precedent, forged authority claims, instruction wording, and tool friction to observe differing adjudication policies. The authors propose a conflict adjudication framework that maps information to interpretation to action, arguing it better captures agent alignment under competing pressures.
By Ivy Zhang
arXiv:2609.21423v1 Announce Type: new
Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...
By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)
The paper argues that deploying generative AI agents requires more than isolated task success; they must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows. The authors introduce two complementary evaluation aspects—operational resilience and considerate participation—to assess how agents recover from blocked work, communicate limits, and adapt to affected people and role boundaries. Using 120 simulated healthcare trajectories across two AI models and twelve stakeholder-derived tasks under varying challenge levels, the study finds that agents shift toward greater human dependence and increased workload as challenge accumulates, while also broadening from task-focused adaptation to task reframing and wider coordination.
By Yuanchen Bai, Zijian Ding, Angelique Taylor
arXiv:2607. 28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system.
By Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru, Darvin Yi, Aakash Sabharwal, Yunzhong He