arXiv AI

Available but Unclaimed: An Empirical Study of Human-AI Synergy

arXiv AI
Aug 19

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

The study investigates how different AI support formats influence human decision-making across two tasks: abstract visual reasoning with RAVEN matrices and deductive logical reasoning with LSAT problems. Findings reveal that in visual reasoning, predictions alone and predicted probabilities best support accuracy and error recovery, while in logical reasoning, LLM explanations outperform other supports. The results suggest that effective human–AI collaboration requires task‑specific support strategies rather than a one‑size‑fits‑all approach.

By Ruth Cohen, Lu Feng, Ayala Bloch, Sarit Kraus
arXiv AI
Sep 18

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

The paper introduces NeuroCognition, a benchmark based on three neuropsychological tests—Raven's Progressive Matrices, Spatial Working Memory, and the Wisconsin Card Sorting Test—to evaluate foundational cognitive abilities in large language models (LLMs). It finds that while LLMs excel on text tasks, their performance drops on image-based and more complex tasks, and they fail different parts of the same tasks compared to humans. NeuroCognition correlates with standard general-capability benchmarks yet measures distinct cognitive skills, highlighting where LLMs align with or diverge from human-like intelligence.

By Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv AI
Jul 24

AI Assistants Overassist

arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.

By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv AI
Sep 16

Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

The paper investigates how to help users monitor their own and an AI system’s competence when using AI assistance. It identifies 30 interventions from experts and organizes them into a design space based on timing, target competence, and source of cue. A large experiment shows that reliability cards and contrasting replies reduce estimation error and overconfidence, though they do not improve task performance.

By Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes, Albrecht Schmidt, Verena Distler, Robin Welsch
arXiv AI
2d ago

When the AI Leaves the Tailorshop: Measuring What an LLM Advisor Leaves Behind in Complex Problem Solving

The study investigates how large language model (LLM) advisors affect complex problem‑solving in a simulated clothing‑factory setting. Two preregistered experiments (N=200 and N=198) found that participants with AI support reported higher confidence and understanding, expended less effort, and in some cases achieved better performance or avoided bankruptcy. Within the AI‑supported group, more frequent changes to the AI’s recommendations were linked to improved unaided performance and knowledge.

By Robin Welsch
arXiv AI
Sep 21

How do LLMs Compute Verbal Confidence

arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...

By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c