arXiv Machine Learning By Anagha Tiwari, Alexander G. Gray, Nick Feamster, Brian Jabarian, Alex Imas, Alex Kale

GUIDE: Generative Utility Inference and Decision Engine

Read the original on arXiv Machine Learning →

GUIDE is a large‑language‑model driven architecture that elicits and infers human user preferences through conversational Bayesian adaptive sampling and symbolic rule‑based learning. It extends adaptive sampling to a wide range of elicitation questions via a flexible type system and initializes domain‑specific preference models using symbolic representations of world knowledge. In simulated investment portfolio optimization, GUIDE outperforms prior methods, LLM‑only baselines, and its own ablated variants by improving cold‑start performance and reducing recommendation regret during early interactions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

Gradient-Aligned Pair Selection for Personalized Preference Optimization

The paper introduces GAP-DPO, a method for personalizing large language models by selecting preference pairs based on gradient alignment with user utility. It formalizes personalized preference learning as a geometry‑aligned optimization problem, showing that off‑policy sampling can shift DPO updates from error correction to reinforcement when preference margins align with utility gradients. Experiments demonstrate that GAP‑DPO improves stylistic fidelity, preference alignment, and overall generation quality over standard DPO variants.

By Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou
arXiv AI
Sep 25

From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs

The paper introduces BaCVA, a Bayesian Context-aware personalized Value Alignment method for large language models. It treats personal values as priors and context-dependent preferences as posteriors, estimating contextual value salience from normative responses and using a dual-view personalization module to infer posterior preferences. Experiments show BaCVA outperforms strong baselines, offering more accurate and data‑efficient personalized value alignment.

By Hanze Guo, Aixuan Song, Jing Yao, Xiangxu Zhang, Xiaoyuan Yi, Xing Xie, Xiao Zhou
arXiv AI
Aug 19

CARA: Cognitive Adaptive Recommendation Agent

CAR A is a recommendation framework that treats recommendation as a structured decision‑making process. It separates recommendation into two stages: candidate filtering, which narrows the search space using coarse preference constraints, and dual‑perspective decision modeling, which captures decisions through affective and rational judgments. A boundary‑aware KTO strategy is introduced to prioritize instructions that the model can solve occasionally but not consistently, thereby enriching preference signals. Experiments on three Amazon Reviews domains show CAR A outperforms baselines, achieving up to a 10.15% relative improvement on most metrics.

By Weijun Gao, Jinyang Dong, Chuanru Ren, Hengxiao Li