arXiv Machine Learning

Conformity Breaks Conformal Prediction

A new study shows that conformal certificates can become invalid when a large language model (LLM) is influenced by peers who unanimously provide a wrong answer, even though the question itself remains unchanged. This phenomenon, termed a score‑mechanism shift, reveals that a model’s calibration for single‑agent scoring does not hold in multi‑agent settings, leading to a drop in coverage from 90% to 74% under unanimous‑wrong peers. The shift also allows attackers to target low‑confidence items, nearly halving coverage for that subgroup while keeping overall averages deceptively high, and can cause systems to act confidently on incorrect answers. whyItMatters":"The findings expose a critical vulnerability in conformal prediction for multi‑agent LLM systems, undermining their reliability and safety in real‑world applications."

arXiv Machine Learning
Sep 17

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

The paper investigates how multi‑agent large language models (LLMs) can correct each other’s mistakes, but also how peer pressure can overturn correct answers. It argues that a safeguard— a ‘brake’ that blocks harmful revisions while allowing beneficial ones— is essentially a correctness probe, and that models’ self‑knowledge (measured by AUROC 0.64–0.89) limits the effectiveness of such a brake. The authors find that even white‑box steering cannot break this ceiling, and that adding information before revision, rather than filtering after, is the more promising approach.

By Yibo Hu
arXiv Machine Learning
3d ago

Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction

This paper introduces a Byzantine‑robust federated retrieval‑augmented generation (RAG) framework that uses aligned calibration and fixed‑membership conformal prediction to ensure that the answer set contains the correct answer with a chosen probability, even when some nodes are compromised. By having all nodes score the same calibration questions and retaining only candidates that could be kept by a plausible group of honest nodes, the method guarantees correctness in finite samples and produces smaller answer sets than simpler approaches. Experiments on medical exam question‑answering tasks with language‑model nodes demonstrate that the method meets the target coverage whenever the number of misbehaving nodes does not exceed the declared bound, while plain averaging often fails.

By Prasanjit Dubey, Aritra Guha, Xiaoming Huo
arXiv Machine Learning
Aug 26

Replicable Conformal Prediction

arXiv:2608.23638v1 Announce Type: cross Abstract: Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration...

By Marios Papamichalis, Regina Ruane, Theofanis Papamichalis