KTO: Model Alignment as Prospect Theoretic Optimization
Read the original on arXiv AI →The paper introduces KTO, a new loss function that aligns large language models with human preferences by directly maximizing a prospect-theoretic utility function rather than log-likelihood of preferences. It builds on the idea that existing alignment objectives implicitly use human-aware losses (HALOs) that reflect biases such as loss aversion. KTO achieves performance comparable to or better than preference-based methods across 1B to 30B parameter models while only requiring a binary desirability signal, and it highlights that no single HALO is universally best, emphasizing the importance of choosing the right inductive biases for a given setting.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.