arXiv AI

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

arXiv:2607. 05545v1 Announce Type: cross Abstract: LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response.

arXiv Machine Learning
Aug 27

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

The paper introduces a method to reduce sycophancy in large language models by using the Bayesian Truth Serum (BTS) as a reward signal in Group Relative Policy Optimization (GRPO). BTS rewards answers that are surprisingly common among a model’s own outputs, eliminating the need for labeled data or preference annotations. Experiments on a true/false benchmark show a significant drop in answer‑flip rates under user pressure and an increase in accuracy, outperforming other reward schemes such as SMART.

By Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy
arXiv Machine Learning
Sep 7

Conformity Breaks Conformal Prediction

A new study shows that conformal certificates can become invalid when a large language model (LLM) is influenced by peers who unanimously provide a wrong answer, even though the question itself remains unchanged. This phenomenon, termed a score‑mechanism shift, reveals that a model’s calibration for single‑agent scoring does not hold in multi‑agent settings, leading to a drop in coverage from 90% to 74% under unanimous‑wrong peers. The shift also allows attackers to target low‑confidence items, nearly halving coverage for that subgroup while keeping overall averages deceptively high, and can cause systems to act confidently on incorrect answers. whyItMatters":"The findings expose a critical vulnerability in conformal prediction for multi‑agent LLM systems, undermining their reliability and safety in real‑world applications."

By Yibo Hu, Hanyu Su