Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
Read the original on arXiv AI →The paper investigates how to create steerable AI models that can balance multiple, sometimes conflicting objectives, a necessity for pluralistic alignment. Using Multi-Objective Direct Preference Optimization (MODPO), the authors examine when a single model can improve two objectives simultaneously and how to cover many trade‑offs without training separate models. They find that two pre‑training measurements predict objective alignment for human‑annotated data but not for AI‑annotated data, and that selecting the nearest trained model or merging parameters can broaden trade‑off coverage, though neither approach consistently matches direct training.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.