arXiv AI By Daniel Commey

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

Read the original on arXiv AI →

arXiv:2601. 22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

The paper introduces PRISMA-LLM, a reporting framework for AI-assisted systematic reviews. It is based on an analysis of 888 review-automation papers, showing a shift toward LLM- and software-driven workflows and inconsistent reporting of evaluation and limitations. The framework separates implementation details from consequence-sensitive evaluation and limitation reporting.

By Miguel Zabaleta, Baihan Lin
arXiv AI
Aug 11

$A^2E$ : An End-to-End Agent Auditing Engine

arXiv:2608. 07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
arXiv AI
Aug 10

An End-to-End Agent Auditing Engine

arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.

By Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou