arXiv AI By Mohammed Bousmah

LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators

Read the original on arXiv AI →

arXiv:2606. 29437v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key question: how can we evaluate not only AI-assisted outputs, but also the interaction process that produced them?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

The paper introduces CoSLR, a Human‑AI collaborative system for systematic literature reviews that incorporates mandatory human checkpoints within a three‑phase pipeline using large language models and Retrieval‑Augmented Generation. In a survey of 63 participants, 42.9 % rated the system’s usability highly, yet 34.9 % indicated they would trust AI‑generated summaries without further human verification after brief interaction. The study highlights that effective human oversight in AI‑assisted literature reviews depends on users’ willingness to engage with the checkpoints, underscoring a calibration issue that interface design must directly address.

By MD Aidul Islam, Malik Abdul Sami, Muhammad Waseem, Zeeshan Rasheed, Kai-kristian Kemell, Zheying Zhang, Pekka Abrahamsson