arXiv Machine Learning

A Distribution Mapping Approach to Counterfactually Fair Reinforcement Learning

arXiv:2608. 08743v1 Announce Type: cross Abstract: Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time.

Hugging Face Trending Papers
Jun 4

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks with observable counterfactual outcomes. Existing datasets either rely on real-world observations without ground-truth counterfactuals or on simplified simulations that fail to capture complex causal dynamics.

Hugging Face Trending Papers
6d ago

Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

The paper introduces Q-learning Penalized Transformer (QPT), a training–inference consistent framework for safe offline reinforcement learning. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost while incorporating a Q-shaped penalty to balance safety, reward maximization, and behavior regularization. The method consistently outperforms strong baselines on 38 DSRL benchmark tasks and adapts robustly to varying constraint thresholds.