arXiv AI

Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift

arXiv:2607. 17384v1 Announce Type: new Abstract: This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles.

arXiv Computer Vision
Aug 27

HEDGE: A Calibrated Ensemble for A/H Recognition

arXiv:2607.12176v2 Announce Type: replace Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the...

By Josep Cabacas-Maso, Ismael Benito-Altamirano, Carles Ventura
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv Machine Learning
5d ago

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

The paper audits a commercial non‑generative System‑1 model on biosecurity‑relevant benchmarks, evaluating accuracy, calibration, error detection, selective prediction, and sensitivity to answer‑option order. It finds that the model’s accuracy varies strongly by task, is reasonably well calibrated when the vendor’s uncertainty field is interpreted correctly, and that option order can cause significant prediction changes—averaging across rotations improves accuracy. The study also shows that applying averaging only to low‑confidence items recovers most of the gain at a lower cost.

By Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
arXiv AI
2d ago

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.

By Amit Singh Bhatti, Vishal Vaddina