The paper investigates how large language models (LLMs) can perform multi-coder qualitative coding by independently coding, debating, and reconciling disagreements. It quantifies the effectiveness of this approach across diverse datasets, identifying key factors—such as codebook length, data similarity, and agent disagreement—that influence coding accuracy. The study finds that intense, unresolved debates improve accuracy but that LLMs still lack adaptive responsiveness to context, leading to design recommendations for automated coding systems.
By Jeongyeon Kim, John Mitchell
arXiv:2608.22417v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive con...
By Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckstr\"om
arXiv:2607. 28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate.
By Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun
arXiv:2608.30543v1 Announce Type: new
Abstract: Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, but existing applications often provide limited methodological t...
By Nadia Jul Jeldtoft, Tariq Yousef
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
By Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao, Zachary Zhang, Yulia L\'apicus, Thomas Han, Kevin He, Min Sun
arXiv:2606. 10881v1 Announce Type: new Abstract: Learner agency and autonomy are foundational to personal development, yet a pervasive "jingle-jangle" fallacy (i.
By Fei Qin, Xiaobo Liu, Yaowen Zhang, Xuming Li, Fei Wang, Mutlu Cukurova, Jingjing Chen, Yu Zhang
The paper introduces Agent-as-Peer-Debriefing, a multi‑agent framework that incorporates peer debriefing into qualitative data analysis with large language models. A Hierarchical Coding Agent generates codes and reflections, which are then refined by three Peer‑Debriefing Agents applying Theory‑Driven, Data‑Driven, or Applied perspectives. Experiments on three datasets show that perspective‑based refinement aligns more closely with human codes than a single‑LLM baseline, and that the choice of perspective offers meaningful trade‑offs.
By Zhimin Lin, Kun Cheng, Zhiyao Shu, Junhua Fang, Juntao Li, Fan Bai, Jie Gao
arXiv:2509.05346v3 Announce Type: replace
Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how thei...
By Bo Yuan, Jiazi Hu
arXiv:2606. 26541v1 Announce Type: new Abstract: Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need.
By Jerome Marston, Tino Kreutzer, Salom\'e Garnier, Ella Boone, Phuong N Pham, Patrick Vinck
Learner agency and autonomy are foundational to personal development, yet a pervasive "jingle-jangle" fallacy (i. e.
arXiv:2505. 00100v2 Announce Type: replace-cross Abstract: Background and Context.
By Ethan Dickey, Andres Bejarano, Rhianna Kuperus, B\'arbara Fagundes
This scoping review examines 421 studies (2015‑2026) on natural language processing applied to student evaluation of teaching comments. It maps the technical evolution from lexicons and classifiers to transformers and large language models, and evaluates four value dimensions. The review identifies a significant gap between actionable outputs (61.3%) and intended‑user evaluation (11.6%), highlighting limited progress in educational value and robustness.
By Jeff Eicher, Rafael da Silva