arXiv:2606. 15559v1 Announce Type: cross Abstract: The transition toward software-defined vehicles concentrates an increasing share of vehicle functionality into distributed software services, where failures propagate through service dependencies and the surface symptom is often several causal hops away from the underlying defect.
By Matthias Wei{\ss}, Athreya Hosahalli Prakash, Falk Dettinger, Nasser Jazdi, Michael Weyrich
arXiv:2602. 01135v3 Announce Type: replace Abstract: Autoregressive models trained via next-token prediction implicitly learn the conditional independence structure of their data-generating process.
By Hugo Math, Rainer Lienhart
arXiv:2606. 09892v1 Announce Type: new Abstract: Textual event records, such as alarm logs, have become an increasingly common data source in engineering and manufacturing systems.
By Xiaofeng Xiao, Jianhong Chen, Qiuzhuang Sun, Naichen Shi, Xubo Yue
arXiv:2607. 22385v1 Announce Type: cross Abstract: Diagnosing the root cause of anomalies is essential for safe industrial operation.
By Amaury Wei, Olga Fink
arXiv:2604. 17616v3 Announce Type: replace Abstract: Root cause analysis (RCA) for time-series anomaly detection is critical for the reliable operation of complex real-world systems.
By Shashank Mishra, Karan Patil, Cedric Schockaert, Didier Stricker, Jason Rambach
arXiv:2608. 01975v1 Announce Type: cross Abstract: Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the inference engine, Python/C++ backend, host CUDA APIs, GPU kernels, and distributed communication.
By Ruilin Xu, Junyi Li, Pengfei Chen, Zongxuan Xie
arXiv:2606. 03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks.
By Taiyu Zhu, Yifan Wu, Weilin Jin, Ying Li, Gang Huang
arXiv:2607. 13548v1 Announce Type: new Abstract: Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches.
By Athira Gopal, Ashwanth Krishnan
arXiv:2605. 22779v2 Announce Type: replace-cross Abstract: Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible.
By Huanchi Wang, Zihang Huang, Yifang Tian, Kristina Dzeparoska, Hans-Arno Jacobsen, Alberto Leon-Garcia
arXiv:2412. 11800v4 Announce Type: replace Abstract: Extracting anomaly causality facilitates diagnostics once monitoring systems detect system faults.
By Mulugeta Weldezgina Asres, Christian Walter Omlin, The CMS-HCAL Collaboration
arXiv:2602. 13807v2 Announce Type: replace Abstract: Time series anomaly detection is critical in many real-world applications, where effective solutions must localize anomalous regions and support reliable decision-making under complex settings.
By Xiaoyu Tao, Yuchong Wu, Mingyue Cheng, Ze Guo, Tian Gao
Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods.