arXiv Machine Learning By Ivan Kruzhilov

Disentangling Optimization Scale from Preference Scale in DPO

Read the original on arXiv Machine Learning →

The paper investigates the Direct Preference Optimization (DPO) objective used for aligning language models, revealing that its coefficient β simultaneously controls both the inverse preference-noise scale and the optimization dynamics. This entanglement causes non‑monotonic policy deviation with respect to β and makes loss values incomparable across different β settings. The authors propose a centered‑softplus reformulation that decouples these effects, allowing independent tuning of the noise scale and learning‑rate, and provides a smooth β←0 limit that reduces to a linear preference‑margin objective.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.