arXiv AI

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

arXiv:2607. 28889v1 Announce Type: cross Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants.

arXiv AI
Sep 12

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

The paper investigates how large language models (LLMs) can perform multi-coder qualitative coding by independently coding, debating, and reconciling disagreements. It quantifies the effectiveness of this approach across diverse datasets, identifying key factors—such as codebook length, data similarity, and agent disagreement—that influence coding accuracy. The study finds that intense, unresolved debates improve accuracy but that LLMs still lack adaptive responsiveness to context, leading to design recommendations for automated coding systems.

By Jeongyeon Kim, John Mitchell
arXiv AI
Aug 3

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

arXiv:2607. 28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate.

By Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun
arXiv AI
Sep 1

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

The paper introduces Agent-as-Peer-Debriefing, a multi‑agent framework that incorporates peer debriefing into qualitative data analysis with large language models. A Hierarchical Coding Agent generates codes and reflections, which are then refined by three Peer‑Debriefing Agents applying Theory‑Driven, Data‑Driven, or Applied perspectives. Experiments on three datasets show that perspective‑based refinement aligns more closely with human codes than a single‑LLM baseline, and that the choice of perspective offers meaningful trade‑offs.

By Zhimin Lin, Kun Cheng, Zhiyao Shu, Junhua Fang, Juntao Li, Fan Bai, Jie Gao
arXiv Machine Learning
Jun 26

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

arXiv:2606. 26541v1 Announce Type: new Abstract: Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need.

By Jerome Marston, Tino Kreutzer, Salom\'e Garnier, Ella Boone, Phuong N Pham, Patrick Vinck
arXiv Machine Learning
Sep 24

From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026

This scoping review examines 421 studies (2015‑2026) on natural language processing applied to student evaluation of teaching comments. It maps the technical evolution from lexicons and classifiers to transformers and large language models, and evaluates four value dimensions. The review identifies a significant gap between actionable outputs (61.3%) and intended‑user evaluation (11.6%), highlighting limited progress in educational value and robustness.

By Jeff Eicher, Rafael da Silva