arXiv Machine Learning

GUIDE: Generative Utility Inference and Decision Engine

GUIDE is a large‑language‑model driven architecture that elicits and infers human user preferences through conversational Bayesian adaptive sampling and symbolic rule‑based learning. It extends adaptive sampling to a wide range of elicitation questions via a flexible type system and initializes domain‑specific preference models using symbolic representations of world knowledge. In simulated investment portfolio optimization, GUIDE outperforms prior methods, LLM‑only baselines, and its own ablated variants by improving cold‑start performance and reducing recommendation regret during early interactions.

arXiv AI
2d ago

Gradient-Aligned Pair Selection for Personalized Preference Optimization

The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.

By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv AI
Sep 25

From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs

The paper introduces BaCVA, a Bayesian Context-aware personalized Value Alignment method for large language models. It treats personal values as priors and context-dependent preferences as posteriors, estimating contextual value salience from normative responses and using a dual-view personalization module to infer posterior preferences. Experiments show BaCVA outperforms strong baselines, offering more accurate and data‑efficient personalized value alignment.

By Hanze Guo, Aixuan Song, Jing Yao, Xiangxu Zhang, Xiaoyuan Yi, Xing Xie, Xiao Zhou
arXiv AI
Aug 19

CARA: Cognitive Adaptive Recommendation Agent

CAR A is a recommendation framework that treats recommendation as a structured decision‑making process. It separates recommendation into two stages: candidate filtering, which narrows the search space using coarse preference constraints, and dual‑perspective decision modeling, which captures decisions through affective and rational judgments. A boundary‑aware KTO strategy is introduced to prioritize instructions that the model can solve occasionally but not consistently, thereby enriching preference signals. Experiments on three Amazon Reviews domains show CAR A outperforms baselines, achieving up to a 10.15% relative improvement on most metrics.

By Weijun Gao, Jinyang Dong, Chuanru Ren, Hengxiao Li
arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv AI
Jul 23

Personalized Recommendation Tool Learning via Autonomous Language Agents

arXiv:2607. 19739v1 Announce Type: cross Abstract: Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks.

By Mingdai Yang, Zhiwei Liu, Weizhi Zhang, Yibo Wang, Hao Peng, Philip Yu
arXiv AI
Sep 17

Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

Re2A is a new framework for situated conversational recommendation that models user interactions within shared physical environments. It introduces rubric-based preference reasoning to explicitly capture user preferences from dialogue history and scene context, and a preference-conditioned optimization to align generated responses with both user satisfaction and situational consistency. Experiments on two SCR datasets show that Re2A outperforms existing methods, providing more precise and context-aware recommendations.

By Dongding Lin, Jian Wang, Xiaoyan Zhao, Wenjie Li
arXiv AI
Sep 25

Learning Better Reasoning for Generative Recommendation with Semantic IDs

The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. It first aligns SIDs with textual and behavioral contexts, then selects candidate reasoning traces that improve ground‑truth item prediction, and finally refines the reasoning policy via reinforcement learning with catalog‑constrained generation and ranking‑aware feedback. Experiments on Amazon Review datasets show Evo‑Rec consistently outperforms existing discriminative, generative, and reasoning‑enhanced recommenders across all metrics.

By Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao