arXiv AI

When the AI Leaves the Tailorshop: Measuring What an LLM Advisor Leaves Behind in Complex Problem Solving

The study investigates how large language model (LLM) advisors affect complex problem‑solving in a simulated clothing‑factory setting. Two preregistered experiments (N=200 and N=198) found that participants with AI support reported higher confidence and understanding, expended less effort, and in some cases achieved better performance or avoided bankruptcy. Within the AI‑supported group, more frequent changes to the AI’s recommendations were linked to improved unaided performance and knowledge.

arXiv AI
Sep 16

Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

The paper investigates how to help users monitor their own and an AI system’s competence when using AI assistance. It identifies 30 interventions from experts and organizes them into a design space based on timing, target competence, and source of cue. A large experiment shows that reliability cards and contrasting replies reduce estimation error and overconfidence, though they do not improve task performance.

By Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch
arXiv AI
Jul 15

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

arXiv:2512. 01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.

By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
arXiv AI
Aug 26

What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

The article examines how AI‑assisted item generation is filtered by a computational evaluator before expert review, focusing on representation, structural screening, and candidate‑form dependence. Through two in‑silico studies of 32,000 Big Five items, the authors show that subtle differences in semantic representation and structural evaluation lead to divergent item selections, even when overall content coverage appears stable. The findings reveal that the evaluator, often treated as a neutral technical step, actually shapes the evidence and wording that psychometricians ultimately review, highlighting its role as a revisable component of measurement design.

By Christopher Brooks (School of Information, University of Michigan)