The paper introduces a method for learning risk scores that remain reliable even when historical data contain unobserved confounders. By treating propensity weights as uncertain and applying sensitivity analysis with Wasserstein distributionally robust optimization, the authors formulate a robust learning problem solvable via an exponential cone program. Experiments on semi‑synthetic UCI data show the approach improves calibration by up to 29.2% over traditional benchmarks and 11.1% over the state of the art, without harming other performance metrics.
By Ryan Edmonds, Yingxiao Ye, Sina Aghaei, Andr\'es G\'omez, \c{C}a\u{g}{\i}l Ko\c{c}yi\u{g}it, Phebe Vayanos
Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios. In this paper, we propose the Cross-Head Attention Uplift Network (CHAUN) and Robust Adversarial Inverse Propensity Score (RA-IPS) method to address these limitations.
arXiv:2606. 27114v1 Announce Type: new Abstract: Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios.
By Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, Feng Zhou
arXiv:2607. 03425v1 Announce Type: new Abstract: Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisions, in potentially high-stakes scenarios.
By Denise Tampieri, Giovanni De Toni, Paolo Giudici
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
By Kosuke Iguchi, Ren Kishimoto
This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.
By Noah Mami\'e, Laurin van den Bergh