arXiv Machine Learning
1d ago

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

The paper investigates how sentence‑specificity scores can guide the selection of revisions in collaborative technical documentation. It compares two predictors—SpeciTeller and a target‑adapted model by Ko et al.—across Wikipedia and three technical corpora, finding that the predictors rank sentences differently and that SpeciTeller can improve direction‑valid selection rates in certain datasets. The study also shows that filtering and token‑length adjustments alter but do not reconcile these ranking differences.

By Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias
arXiv Computation and Language
Sep 22

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

The paper introduces CoSLR, a Human‑AI collaborative system for systematic literature reviews that incorporates mandatory human checkpoints within a three‑phase pipeline using large language models and Retrieval‑Augmented Generation. In a survey of 63 participants, 42.9 % rated the system’s usability highly, yet 34.9 % indicated they would trust AI‑generated summaries without further human verification after brief interaction. The study highlights that effective human oversight in AI‑assisted literature reviews depends on users’ willingness to engage with the checkpoints, underscoring a calibration issue that interface design must directly address.

By MD Aidul Islam, Malik Abdul Sami, Muhammad Waseem, Zeeshan Rasheed, Kai-kristian Kemell, Zheying Zhang, Pekka Abrahamsson