The paper introduces a method for learning risk scores that remain reliable even when historical data contain unobserved confounders. By treating propensity weights as uncertain and applying sensitivity analysis with Wasserstein distributionally robust optimization, the authors formulate a robust learning problem solvable via an exponential cone program. Experiments on semi‑synthetic UCI data show the approach improves calibration by up to 29.2% over traditional benchmarks and 11.1% over the state of the art, without harming other performance metrics.
By Ryan Edmonds, Yingxiao Ye, Sina Aghaei, Andr\'es G\'omez, \c{C}a\u{g}{\i}l Ko\c{c}yi\u{g}it, Phebe Vayanos
Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios. In this paper, we propose the Cross-Head Attention Uplift Network (CHAUN) and Robust Adversarial Inverse Propensity Score (RA-IPS) method to address these limitations.
arXiv:2606. 27114v1 Announce Type: new Abstract: Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios.
By Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, Feng Zhou
arXiv:2607. 03425v1 Announce Type: new Abstract: Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisions, in potentially high-stakes scenarios.
By Denise Tampieri, Giovanni De Toni, Paolo Giudici
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
By Kosuke Iguchi, Ren Kishimoto
This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.
By Noah Mami\'e, Laurin van den Bergh
arXiv:2608. 12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions.
By Xuefei Wang, Jun Han, Zixuan Wang, Qingkai Zeng, Xiao Wang, Ruijie Wang, Jianxin Li
The paper proposes an incremental recommendation system that uses causal modeling to avoid delivering redundant recommendations. By leveraging existing holdback data and a dual‑threshold targeting policy, the authors reduce recommendation impressions by 7% without harming overall content consumption. Joint training with holdback data also improves the calibration of the treated model, suggesting better generalisable representations than purely observational models.
By Athanasios Vlontzos, David Gustafsson, Michael O'Riordan, Ciar\'an M. Gilligan-Lee
The paper proposes an incremental recommendation approach that uses a causal model built from existing holdback data to avoid delivering redundant recommendations. By applying a dual‑threshold targeting policy, the system only recommends content when the likelihood of a treated stream is high and the likelihood of an organic stream is low, thereby reducing recommendation impressions by 7% without hurting overall consumption. Joint training with holdback data also improves the calibration of the treated head, suggesting that causal models capture more generalisable representations than purely observational models.
The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.
The paper introduces a systematic benchmark for evaluating explainable methods that attribute temporal interactions in sequential recommendation systems. Using a dual-model masking metric, it assesses ten XAI techniques across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens datasets, revealing that gradient-based methods like GradientSHAP and Integrated Gradients are the most faithful and robust. It also finds that raw attention weights are unreliable, while gradient-weighted attention works better on short sequences but degrades on longer horizons, and that faithful methods capture genuine task structure rather than recency or popularity bias.
By Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.
By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun