arXiv:2607. 28889v1 Announce Type: cross Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants.
By Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He
arXiv:2609.16487v1 Announce Type: new
Abstract: We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid...
By Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal, Aditya Bansal, Rui Wang, Charles Menguy, Swati Jain
arXiv:2606. 24839v1 Announce Type: new Abstract: Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics.
By Tian Zheng, Kai-Tai Hsu
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
The paper investigates how large language models (LLMs) can perform multi-coder qualitative coding by independently coding, debating, and reconciling disagreements. It quantifies the effectiveness of this approach across diverse datasets, identifying key factors—such as codebook length, data similarity, and agent disagreement—that influence coding accuracy. The study finds that intense, unresolved debates improve accuracy but that LLMs still lack adaptive responsiveness to context, leading to design recommendations for automated coding systems.
By Jeongyeon Kim, John Mitchell
arXiv:2609.16793v1 Announce Type: cross
Abstract: People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a bet...
By Robin Welsch, Michelle Rausch, Pascal Knierim, Thomas Kosch, Jochen Kuhn, Albrecht Schmidt, Daniela Fernandes
The paper compares human group discussions with large language model (LLM) deliberation traces on various reasoning tasks, finding that both humans and LLMs exhibit an assembly bonus asymmetry where discussion benefits the average member more than the best initial member. While LLM groups mirror some outcome-level patterns of human deliberation, they differ in process-level behaviors: they tend to follow majorities, surface less unique information, and converge earlier. Interventions inspired by human group‑decision research yield modest outcome improvements but do not eliminate coordination bottlenecks.
By Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch
arXiv:2608.22417v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive con...
By Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckstr\"om
arXiv:2608.30373v1 Announce Type: new
Abstract: Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. Howev...
By Minsoo Song, Chanwoo Kim, Sugyeong Eo, Chanjun Park
The study investigates how different AI support formats influence human decision-making across two tasks: abstract visual reasoning with RAVEN matrices and deductive logical reasoning with LSAT problems. Findings reveal that in visual reasoning, predictions alone and predicted probabilities best support accuracy and error recovery, while in logical reasoning, LLM explanations outperform other supports. The results suggest that effective human–AI collaboration requires task‑specific support strategies rather than a one‑size‑fits‑all approach.
By Ruth Cohen, Lu Feng, Ayala Bloch, Sarit Kraus
arXiv:2605. 29928v2 Announce Type: replace-cross Abstract: As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downstream consequences for moderation, evaluation, and decision-making.
By Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong, Ting-Hao 'Kenneth' Huang, Dongwon Lee
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang