arXiv AI By Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Read the original on arXiv AI →

arXiv:2509. 03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

CroCo introduces cross‑lingual contrastive preference tuning on self‑generations, extending prior English‑only methods to 14 high‑ and low‑resource languages. A reward model trained solely on English preferences, applied to a multilingual base, yields effective within‑language rankings and improves performance in both monolingual and multilingual settings without catastrophic forgetting. The approach requires on‑policy data; off‑policy responses and online preference optimization offer limited gains, yet on structured tasks CroCo matches or surpasses the base model in most languages, and on open‑ended generation it wins 28/30 judge evaluations across 15 languages.

By Mike Zhang, Ali Basirat, Desmond Elliott