Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
Read the original on arXiv AI →The study examined how four large language models (GPT‑5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) scored 18 simulated Japanese‑language AI‑to‑AI counseling sessions compared to ratings from 15 human counseling experts. Each model evaluated every transcript three times on four motivational‑interviewing‑informed dimensions and overall quality, consistently giving higher scores for softening sustain talk and overall quality than the expert panel, though the magnitude varied by model. Run‑to‑run reliability (intraclass correlation coefficients ranging from .33 to .96) did not predict closer alignment with expert judgments, and the models’ ability to discriminate counselor conditions was distinct from both reliability and alignment.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.