arXiv:2608. 12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images.
By Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin, Abhishek Mukherji
arXiv:2606. 09852v1 Announce Type: cross Abstract: High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainability are essential.
By Ikbel Ghrab, Mohamed Dhieb, Ismail Khenissi, Ines Abdeljaoued-Tej
arXiv:2606. 09090v1 Announce Type: cross Abstract: Developers increasingly provide AI coding assistants with persistent context through configuration files such as CLAUDE.
By Christoph Treude, Sebastian Baltes
arXiv:2602.02475v2 Announce Type: replace
Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy...
By Shraddha Barke, Arnav Goyal, Alind Khare, Avaljot Singh, Suman Nath, Chetan Bansal
arXiv:2607. 24348v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature.
By Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert
The paper examines how Vision Language Models (VLMs) can automatically extract structured procedural knowledge from industrial troubleshooting guides, which are typically flowchart-like diagrams combining spatial layout and technical language. It evaluates two VLMs using two prompting strategies—standard instruction-guided and an augmented approach that highlights layout patterns—and finds that each model shows different trade-offs between sensitivity to layout and robustness to semantic content. These insights help determine which VLM and prompting method is most suitable for integrating such guides into operator support systems.
By Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis
arXiv:2607. 03833v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have achieved remarkable success in Text-to-SQL tasks, their deployment in real-world environments is hindered by latent reliability issues.
By Hanqing Wang, Yongdong Chi, Jian Yang, Lei Yang, Jiehui Zhao, Yun Chen, Guanhua Chen
arXiv:2606. 31567v1 Announce Type: cross Abstract: Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety.
By Shayne Longpre, Elaine Zhu, Carson Ezell, Avijit Ghosh, Sean McGregor, Kevin Paeth, Kevin Klyman, Sayash Kapoor, Rishi Bommasani, Ruth Appel, Gregory Strom, Lauren McIlvenny, Mark M. Jaycox, Peter Slattery, Nathan Butters, Arvind Narayanan, Percy Liang, Alex Pentland
SAGE is a governed multi‑stage LLM pipeline that transforms enterprise guideline documents—containing narrative text, tables, and images—into structured artifacts. It uses a shared versioned rule store, schema‑validated contracts, and provenance tracking to validate, score, and reconcile extracted rules, automatically approving high‑confidence outputs while flagging uncertain items for human review. In a test on 120 documents, SAGE reduced processing time from days to 20–100 minutes and achieved a 96% success rate with only 3.2% hallucination.
By Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli, Srinivasa Karthikeya Reddy Kovvuri, Abhishek Mukherji
The paper presents a systematic analysis of five state‑of‑the‑art automated program repair agents, tracing their decision‑making across 500 real‑world repair tasks. It finds that while the agents perform well on simple fixes, they struggle with logic‑intensive bugs, often producing verbose, overfitted patches that pass tests without addressing root causes. Key bottlenecks identified include poor test generation, limited regression test selection, and reliance on primitive tooling without access to debuggers or advanced program analysis tools.
By Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
arXiv:2608. 10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents.
By You Lu, Kun Zhang, Bihuan Chen, Xin Peng
arXiv:2602. 17990v2 Announce Type: replace Abstract: Multi-agent LLM systems that generate structured workflows from natural-language requests are now deployed in production across cloud automation, DevOps, and enterprise process orchestration.
By Madhav Kanda, Sharad Agarwal, Rodrigo Fonseca, Alok Gautam Kumbhare, Pedro Las-Casas