arXiv AI By Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen

Distributionally Robust Listwise Preference Optimization

Read the original on arXiv AI →

arXiv:2607. 01715v1 Announce Type: new Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.