arXiv:2607. 18127v1 Announce Type: cross Abstract: With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and reliable anomaly detection to maintain service availability.
By Thu T. H. Doan, Mohammad Saiful Islam, Andriy Miranskyy, Ngoc-Thanh Nguyen, Rogardt Heldal, Patrizio Pelliccione
arXiv:2608. 08968v1 Announce Type: cross Abstract: Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graphs.
By Yifang Tian, Yaming Liu, Zichun Chong, Zihang Huang, Yiran Li, Hans-Arno Jacobsen
arXiv:2606. 29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data.
By Yuanhong Cai, Xiaohui Nie, Kanglin Yin, Changhua Pei, Yongqian Sun, Shenglin Zhang, Haibin Liu, Guiyang Liu, Xidao Wen, Fang Situ, Dan Pei
arXiv:2607. 13548v1 Announce Type: new Abstract: Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches.
By Athira Gopal, Ashwanth Krishnan
Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods.
arXiv:2412. 11800v4 Announce Type: replace Abstract: Extracting anomaly causality facilitates diagnostics once monitoring systems detect system faults.
By Mulugeta Weldezgina Asres, Christian Walter Omlin, The CMS-HCAL Collaboration
arXiv:2606. 15559v1 Announce Type: cross Abstract: The transition toward software-defined vehicles concentrates an increasing share of vehicle functionality into distributed software services, where failures propagate through service dependencies and the surface symptom is often several causal hops away from the underlying defect.
By Matthias Wei{\ss}, Athreya Hosahalli Prakash, Falk Dettinger, Nasser Jazdi, Michael Weyrich
arXiv:2509. 06419v2 Announce Type: replace Abstract: Time-series anomaly detection is crucial in AIOps for maintaining large-scale service reliability.
By Xudong Mou, Rui Wang, Tiejun Wang, Zexin Wu, Fangda Guo, Jie Sun, Shiru Chen, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Hao Peng, Chunming Hu, Xudong Liu, Renyu Yang
arXiv:2604. 17616v3 Announce Type: replace Abstract: Root cause analysis (RCA) for time-series anomaly detection is critical for the reliable operation of complex real-world systems.
By Shashank Mishra, Karan Patil, Cedric Schockaert, Didier Stricker, Jason Rambach
arXiv:2602. 13807v2 Announce Type: replace Abstract: Time series anomaly detection is critical in many real-world applications, where effective solutions must localize anomalous regions and support reliable decision-making under complex settings.
By Xiaoyu Tao, Yuchong Wu, Mingyue Cheng, Ze Guo, Tian Gao
arXiv:2607. 27290v1 Announce Type: new Abstract: Modern telecommunication, cloud, and microservice systems emit correlated alarm cascades when components fail.
By Lei Zan, Keli Zhang, Shifeng Xie, Jiale Zheng, Zehao Xiao, Zhiwei Dong, Ke Zhang, Ruichu Cai, Malik Tiomoko, Lujia Pan
arXiv:2603. 25538v3 Announce Type: replace Abstract: Automated incident management is critical for microservice reliability.
By Wenzhuo Qian, Hailiang Zhao, Ziqi Wang, Zhipeng Gao, Jiayi Chen, Zhiwei Ling, Shuiguang Deng