arXiv:2601. 22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes.
By Daniel Commey
arXiv:2608. 10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents.
By You Lu, Kun Zhang, Bihuan Chen, Xin Peng
arXiv:2608.30022v1 Announce Type: new
Abstract: Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing ap...
By Ashvin Gupta, Denys Prociuk, Alessandra Russo, Brendan C. Delaney
The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
By Md Shohel Arman, Igor Molybog
arXiv:2604. 05435v2 Announce Type: replace Abstract: Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions.
By Akshat Dasula, Prasanna Desikan, Jaideep Srivastava, Shivali Dalmia, Abhishek Mukherji
The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.
By Miguel Zabaleta, Baihan Lin