arXiv AI By Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

Read the original on arXiv AI →

arXiv:2607. 28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

The paper investigates how large language models (LLMs) can perform multi-coder qualitative coding by independently coding, debating, and reconciling disagreements. It quantifies the effectiveness of this approach across diverse datasets, identifying key factors—such as codebook length, data similarity, and agent disagreement—that influence coding accuracy. The study finds that intense, unresolved debates improve accuracy but that LLMs still lack adaptive responsiveness to context, leading to design recommendations for automated coding systems.

By Jeongyeon Kim, John Mitchell