The paper investigates how large language models (LLMs) can perform multi-coder qualitative coding by independently coding, debating, and reconciling disagreements. It quantifies the effectiveness of this approach across diverse datasets, identifying key factors—such as codebook length, data similarity, and agent disagreement—that influence coding accuracy. The study finds that intense, unresolved debates improve accuracy but that LLMs still lack adaptive responsiveness to context, leading to design recommendations for automated coding systems.
By Jeongyeon Kim, John Mitchell
arXiv:2608.22417v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive con...
By Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckstr\"om
arXiv:2607. 28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate.
By Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun
arXiv:2608.30543v1 Announce Type: new
Abstract: Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, but existing applications often provide limited methodological t...
By Nadia Jul Jeldtoft, Tariq Yousef
arXiv:2606. 12422v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices.
By Zewei Tian, Alex Liu, Lief Esbenshade, Michael Xiao, Zachary Zhang, Yulia L\'apicus, Thomas Han, Kevin He, Min Sun
arXiv:2606. 10881v1 Announce Type: new Abstract: Learner agency and autonomy are foundational to personal development, yet a pervasive "jingle-jangle" fallacy (i.
By Fei Qin, Xiaobo Liu, Yaowen Zhang, Xuming Li, Fei Wang, Mutlu Cukurova, Jingjing Chen, Yu Zhang