arXiv Machine Learning

Conditionally Identifiable Latent-Environment Modeling for Out-of-Distribution Recommendation

arXiv:2608. 03647v1 Announce Type: cross Abstract: Out-of-distribution (OOD) recommendation is vulnerable to preference shifts induced by a latent environment.

arXiv Machine Learning
Sep 3

GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation

GenCAR introduces a method for out‑of‑distribution recommendation that balances utility and risk by controlling the proxy‑label false discovery rate (FDR). It frames the problem as an α‑Valid Counterfactual Recommendation (α‑VCR) task, coupling counterfactual supervision with calibrated set selection using conformal p‑values and Benjamini–Hochberg filtering. The approach theoretically bounds counterfactual approximation error and guarantees finite‑sample, distribution‑free FDR control under various dependence assumptions, and empirical results show improved OOD candidate recovery across benchmarks.

By Qianqian Wang, Yunshan Li, Jiawen Zeng, Wenwu Gong, Lili Yang
arXiv Machine Learning
Jun 9

The Value of Personalized Recommendations: Evidence from Netflix

arXiv:2511. 07280v5 Announce Type: replace-cross Abstract: Personalized recommendation systems shape much of user choice online, yet their targeted nature makes separating out the value of recommendation and the underlying goods challenging.

By Kevin Zielnicki, Guy Aridor, Aur\'elien Bibaut, Allen Tran, Winston Chou, Nathan Kallus
arXiv Machine Learning
Sep 10

PUID: A Personalized Deconfounding Framework for Recommender Systems under Hidden Confounding

The paper introduces PUID, a Personalized Unobserved-Confounding-aware Interaction Deconfounder, designed to mitigate hidden confounding in recommender systems without relying on costly randomized controlled trials. PUID estimates user-item level sensitivity bounds using an entropy-based method that gauges the strength of hidden confounding from the mutual information between observed features and exposure status. An adversarial optimization strategy and a benchmark-guided variant (BPUID) further enhance robustness and predictive accuracy, and experiments on three real-world datasets show consistent outperformance over state-of-the-art baselines.

By Zongyu Li
arXiv Machine Learning
Sep 22

Multiple latent orderings better predict language model preferences

The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.

By Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young
arXiv AI
Sep 15

Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself

The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.

By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
arXiv AI
Sep 18

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.

By Noah Mami\'e, Laurin van den Bergh