arXiv AI

Reinforcement Learning and Consumption-Savings Behavior

arXiv:2510. 20748v2 Announce Type: replace-cross Abstract: This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns.

arXiv Machine Learning
Jul 7

Understanding electricity consumption behaviour through Inverse Reinforcement Learning

arXiv:2607. 03176v1 Announce Type: new Abstract: Understanding how households consume electricity in response to socioeconomic and climatic drivers is important for decision-makers designing energy policies in a changing climate and under geopolitical tensions.

By Enrico Cofler, Carlos Rodriguez-Pardo, Matteo Giuliani, Andrea Castelletti, Massimo Tavoni
arXiv Machine Learning
Sep 10

Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions

The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.

By Defu Cao, Angela Zhou
arXiv AI
Sep 7

Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

The paper presents a reinforcement learning framework for designing solar PV adoption policies under uncertainty, integrating RL with a stochastic agent‑based model to simulate yearly adoption over a 16‑year horizon. Policymakers can choose annual incentives such as grants, subsidised loans, and feed‑in tariffs, and the study evaluates three RL algorithms—PPO, SAC, and TD3—within a scalarised reward framework that balances adoption gains against costs. Results show clear trade‑off patterns, with TD3 yielding the highest adoption at higher cost, PPO achieving the lowest cost with fewer adopters, and a balanced PPO policy offering a middle ground, all outperforming static baseline policies.

By Iias Faiud, Jonaid Shianifar, Michael Schukat, Karl Mason
arXiv AI
Jul 9

Can Reinforcement Learning Efficiently Discover Price Manipulation?

arXiv:2607. 06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates.

By Ioanna-Yvonni Tsaknaki, Andrea Macr\`i, Fabrizio Lillo