arXiv:2608. 08422v1 Announce Type: cross Abstract: Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback.
By Zhaoyang Shi
GenCAR introduces a method for out‑of‑distribution recommendation that balances utility and risk by controlling the proxy‑label false discovery rate (FDR). It frames the problem as an α‑Valid Counterfactual Recommendation (α‑VCR) task, coupling counterfactual supervision with calibrated set selection using conformal p‑values and Benjamini–Hochberg filtering. The approach theoretically bounds counterfactual approximation error and guarantees finite‑sample, distribution‑free FDR control under various dependence assumptions, and empirical results show improved OOD candidate recovery across benchmarks.
By Qianqian Wang, Yunshan Li, Jiawen Zeng, Wenwu Gong, Lili Yang
arXiv:2511. 07280v5 Announce Type: replace-cross Abstract: Personalized recommendation systems shape much of user choice online, yet their targeted nature makes separating out the value of recommendation and the underlying goods challenging.
By Kevin Zielnicki, Guy Aridor, Aur\'elien Bibaut, Allen Tran, Winston Chou, Nathan Kallus
arXiv:2606. 04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footprint (PCF) are almost never available at catalog scale.
By Noah Lund Syrdal, Anders Vestrum, Jorgen Bergh
The paper introduces PUID, a Personalized Unobserved-Confounding-aware Interaction Deconfounder, designed to mitigate hidden confounding in recommender systems without relying on costly randomized controlled trials. PUID estimates user-item level sensitivity bounds using an entropy-based method that gauges the strength of hidden confounding from the mutual information between observed features and exposure status. An adversarial optimization strategy and a benchmark-guided variant (BPUID) further enhance robustness and predictive accuracy, and experiments on three real-world datasets show consistent outperformance over state-of-the-art baselines.
By Zongyu Li
arXiv:2601. 02322v2 Announce Type: replace-cross Abstract: A common approach to out-of-distribution prediction restricts models to causal or invariant covariates to avoid spurious associations that may change across environments.
By Shuozhi Zuo, Yixin Wang
The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.
By Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.
By Noah Mami\'e, Laurin van den Bergh
arXiv:2608. 14011v1 Announce Type: cross Abstract: Generative recommendation autoregressively generates the semantic IDs of the target item, unifying preference modeling and index retrieval within the shared token space.
By Haokai Ma, Aoqi Hu, Yueao Xing, Ruobing Xie, Yonghui Yang, Teng Tu, Lei Meng, Tat-Seng Chua
arXiv:2606. 03866v1 Announce Type: cross Abstract: Scaling recommender systems via large language models (LLMs) has become a prominent trend in the industry.
By Yuecheng Li, Zeyu Song, Jing Yao, Chi Lu, Peng Jiang, Kun Gai
arXiv:2608.21243v1 Announce Type: cross
Abstract: Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equally informative. Real logs combi...
By Zichun Jin, Zihan Zhou, Yinan Liu, Bin Wang, Xiaochun Yang