arXiv:2606. 03866v1 Announce Type: cross Abstract: Scaling recommender systems via large language models (LLMs) has become a prominent trend in the industry.
By Yuecheng Li, Zeyu Song, Jing Yao, Chi Lu, Peng Jiang, Kun Gai
The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.
By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv:2606. 05828v1 Announce Type: new Abstract: As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm.
By Zeyu Gan, Huayi Tang, Yong Liu
The paper introduces BaCVA, a Bayesian Context-aware personalized Value Alignment method for large language models. It treats personal values as priors and context-dependent preferences as posteriors, estimating contextual value salience from normative responses and using a dual-view personalization module to infer posterior preferences. Experiments show BaCVA outperforms strong baselines, offering more accurate and data‑efficient personalized value alignment.
By Hanze Guo, Aixuan Song, Jing Yao, Xiangxu Zhang, Xiaoyuan Yi, Xing Xie, Xiao Zhou
CAR A is a recommendation framework that treats recommendation as a structured decision‑making process. It separates recommendation into two stages: candidate filtering, which narrows the search space using coarse preference constraints, and dual‑perspective decision modeling, which captures decisions through affective and rational judgments. A boundary‑aware KTO strategy is introduced to prioritize instructions that the model can solve occasionally but not consistently, thereby enriching preference signals. Experiments on three Amazon Reviews domains show CAR A outperforms baselines, achieving up to a 10.15% relative improvement on most metrics.
By Weijun Gao, Jinyang Dong, Chuanru Ren, Hengxiao Li
arXiv:2606. 30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified.
By Irena Saracay, Ludwig Schmidt, Carlos Guestrin