arXiv:2609.16793v1 Announce Type: cross
Abstract: People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a bet...
By Robin Welsch, Michelle Rausch, Pascal Knierim, Thomas Kosch, Jochen Kuhn, Albrecht Schmidt, Daniela Fernandes
Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond.
arXiv:2607. 13562v1 Announce Type: new Abstract: Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question.
By Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro
arXiv:2609.24644v1 Announce Type: cross
Abstract: As people turn to generative AI for financial advice, these systems can personalize how they communicate and what they say. Whether these forms of pe...
By Hasibur Rahman, Benjamin R. Cowan, Smit Desai
Study finds non-experts deferred to LLM-based diagnostic assistance, even when it was wrong, while clinicians caught AI errors.
By Adam Zewe | MIT News
The paper investigates how to help users monitor their own and an AI system’s competence when using AI assistance. It identifies 30 interventions from experts and organizes them into a design space based on timing, target competence, and source of cue. A large experiment shows that reliability cards and contrasting replies reduce estimation error and overconfidence, though they do not improve task performance.
By Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch