arXiv AI By Tony Tu, Sayan Chakraborty, Ruomeng Xu, Tony Qin, Austin Tian

MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

Read the original on arXiv AI →

arXiv:2608. 14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.

By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv Machine Learning
Sep 22

Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time

Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.

By Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot, Amit Dhanda, Aman Chadha, Kapil Wanaskar, Vinija Jain, Amitava Das
arXiv Machine Learning
Jun 16

GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization

arXiv:2606. 16771v1 Announce Type: new Abstract: As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities.

By Haotian Liu, Yihao Liu, Jingwei Ni, Siyuan Huang, Xinpeng Liu, Pengyu Cheng, Jiajun Song, Ruijin Ding, Junfeng Li, Zhechao Yu, Mengyu Zhou, Hongteng Xu, Xiaoxi Jiang, Guanjun Jiang