arXiv:2608. 09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps.
By Yikai Zhao, Pradeep Kumar Misra, Saurabh Pandey
arXiv:2608. 05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows.
By Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang
arXiv:2606. 17203v1 Announce Type: cross Abstract: Multi-agent AI systems are increasingly used to automate software engineering tasks including requirements analysis, architecture design, test generation, and traceability linking.
By Mohamed Essam, Kareem Wael, Azza Hassan, Ahmed Haitham, Mahmoud Soliman, Samer Saber, Ibrahim Habib
arXiv:2607. 02116v1 Announce Type: new Abstract: Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of provenance, version identity, integrity, traceability, or point-in-time reconstruction.
By Misha Sulpovar (PromptOwl, LLC), Benn R. Konsynski (Goizueta Business School, Emory University), Qaish Kanchwala (IBM Research), Gabe Goodhart (IBM Research)
arXiv:2607. 26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible.
By Rwaida Alssadi, Muntaser Syed, Balaji Kasula, Lamine Deen, Majed Alotaibi, Mohammed Alghamdi, Tyler Ton, Ali Alqarni, Marius Silaghi
arXiv:2607. 17242v1 Announce Type: cross Abstract: Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch.
By Md Erfan, Ahmed Ryan, Md Rayhanur Rahman
SIDScope is a diagnostic tool that evaluates Semantic‑ID interfaces used in generative recommendation systems. It normalizes item‑to‑code artifacts, verifies provenance, profiles mapping structure, and compares revisions while tracking path‑to‑item outcomes in generated traces. Using nine tokenizer exports from Amazon and Yelp data, SIDScope shows that interface health depends on multiple signals and reveals gaps in prefix alignment, trace accounting, and mapping refresh effects.
By Jiandong Ding, Huijie Qin, Tiandeng Wu, Yi Cao
SIDScope is a diagnostic tool that evaluates Semantic-ID mappings used between item tokenizers and generative recommenders. It normalizes artifacts, verifies provenance, profiles mapping structure, compares revisions, and tracks path-to-item outcomes in generated traces. Using data from Amazon and Yelp, SIDScope shows that interface health depends on multiple signals, revealing gaps in prefix alignment, trace accounting, and refresh handling that affect model reuse.
arXiv:2606. 04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, environments, and other agents.
By Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Mingkai Zhang, Yanming Zhu
arXiv:2608.28790v1 Announce Type: cross
Abstract: Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from ticket text, historical cases, system logs, an...
By Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao, Huseyin Uzunalioglu
arXiv:2604. 03447v2 Announce Type: replace-cross Abstract: LLM-based software engineering assistants often reason over multiple artifacts, including code, documentation, signatures, and tests, even when those artifacts are incomplete or mutually inconsistent.
By Noshin Ulfat, Ahsanul Ameen Sabit, Soneya Binta Hossain
TRACE is a system designed to bridge the grounding contract gap in LitTraceQA by combining target-aware retrieval, independent typed evidence localization, multimodal table extraction, and schema-driven table construction. It indexes 27,487 papers using multiple representations while preserving question targets, predicts observation units for tables, and assembles rows with evaluator-compatible key normalization. On the official test set, TRACE achieves a 0.760613 overall score, with high paper F1, evidence F1, and multiple-choice accuracy, though table-row and macro cell performance remain lower.
By Sachin Gupta, Divya Godara