Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
Read the original on arXiv AI →The study evaluates how sixteen language models from ten families generate diverse formulations of a psychotherapeutic case, finding an average semantic diversity of 1.69 distinct formulations versus 1.43 for a single-model baseline. It introduces the Vendi Score to quantify diversity and defines a per-model dissent metric to identify the most divergent voice within an ensemble. The analysis shows that model identity significantly influences dissent, but this effect varies across model pairs and panel compositions, indicating that ensemble dispersion is a measurable property rather than an assumed one.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.