arXiv AI By Donghwan Kim

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

Read the original on arXiv AI →

arXiv:2607. 20768v1 Announce Type: cross Abstract: Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

The study evaluates how sixteen language models from ten families generate diverse formulations of a psychotherapeutic case, finding an average semantic diversity of 1.69 distinct formulations versus 1.43 for a single-model baseline. It introduces the Vendi Score to quantify diversity and defines a per-model dissent metric to identify the most divergent voice within an ensemble. The analysis shows that model identity significantly influences dissent, but this effect varies across model pairs and panel compositions, indicating that ensemble dispersion is a measurable property rather than an assumed one.

By Mario Vega-Barbas, Lidia Mora-Valenciano, Iv\'an Pau, Fernando Seoane, Farhad Abtahi
arXiv Computation and Language
Aug 31

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

The paper investigates whether stacking multiple defenses around large language models (LLMs) truly compounds security. Using the Adversary Access‑Tier Model (AATM) and a cost‑tiering system, the authors analyze a seven‑layer defense stack and find that failure correlations between layers are consistently positive, meaning the residual attack success is higher than the multiplicative prediction. Despite high coverage and low false refusals, the stack’s performance is largely driven by common architectural causes rather than diverse, independent defenses.

By Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed