arXiv Machine Learning

Patterning in Practice: Debiasing Reward Models with Susceptibilities

The paper introduces patterning, a reweighting technique that adjusts preference pairs based on their susceptibility to bias, to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. Using this method, the authors achieve a +14.2 ± 1.2 percentage point improvement on the RM‑Bench Hard split while maintaining overall accuracy, matching the best reported Hard‑split gain from a comparable model. The study also demonstrates that the learned weights are interpretable, transferable across Gemma variants, and partially effective on Llama 3.1 8B.

arXiv Machine Learning
Jun 18

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.

By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
arXiv AI
Aug 24

Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation

The paper investigates two strategies for improving large language model (LLM) evaluation: specialized judge weights and rule‑based deferral policies. Experiments on nearly 100,000 rubric‑conditioned samples show that correct rubrics boost accuracy, while incorrect ones hurt it, and that splitting training data into criterion‑specific experts can severely degrade performance unless the experts are warm‑started from a unified model. The authors demonstrate that lightweight deferral cascades can match or exceed the accuracy of larger standalone judges at a fraction of the compute cost, and they provide practical design rules for building efficient, reliable LLM evaluators.

By Ye Chen, Weining Zhang
arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv AI
Aug 25

XTC: Head-Aware Sampling by Excluding Top Choices

XTC (Exclude Top Choices) is a lightweight, head‑aware decoding operator that improves diversity in autoregressive language models by removing overly probable tokens that dominate the next‑token distribution. It works by identifying tokens above a plausibility threshold, probabilistically excluding the dominant choices, and renormalizing the remaining distribution. Across 60 experiments on models such as Gemma 3 and DeepSeek R1, XTC boosts Distinct‑2 scores by 11–15 % and cuts repeat trigrams by 27–47 %, while a Mechanical Turk study shows a 62.3 % preference for XTC‑generated text without loss of fluency.

By Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv