arXiv:2607. 11357v1 Announce Type: new Abstract: Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience.
By Yongqian Sun, Rongchen Gao, Yu Luo, Wenwei Gu, Shenglin Zhang, Qingyi Guo, Qiuai Fu, Yaoliang Wu, Dan Pei
arXiv:2606. 29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data.
By Yuanhong Cai, Xiaohui Nie, Kanglin Yin, Changhua Pei, Yongqian Sun, Shenglin Zhang, Haibin Liu, Guiyang Liu, Xidao Wen, Fang Situ, Dan Pei
arXiv:2609.15161v1 Announce Type: cross
Abstract: Large language model (LLM) driven multi-agent systems have shown promise in complex clinical reasoning, yet existing approaches rely on static strate...
By Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang
The paper introduces AGENTSCOPE, a neuro‑symbolic method for diagnosing failures in large language model agents. It abstracts agent trajectories into structured representations and employs neural invariants to define behavior properties. Using LLM‑guided reasoning on these abstractions, AGENTSCOPE identifies both the failure step and its type, outperforming existing techniques on several datasets.
By Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, Mao Yang
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning.
arXiv:2608.21810v1 Announce Type: cross
Abstract: Clinical decision-making is inherently experience-driven: physicians progressively refine their reasoning by synthesizing patient history, multimodal...
By Md Asaduzzaman Jabin, Khoa Le, Lin Zhao, Tianming Liu