Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning
Read the original on arXiv AI →The paper introduces MI‑SARSA, an on‑policy temporal‑difference algorithm that incorporates mutual‑information regularization to model bounded rationality in reinforcement learning. By penalizing state‑specific deviations from a learned marginal action prior, the algorithm selectively uses state information only when the expected return outweighs the informational cost, yielding a reward‑complexity tradeoff. MI‑SARSA also predicts reaction times, showing that stronger information penalties lead to simpler policies, lower control costs, and faster responses, while regularization mitigates performance loss after environmental shifts at the expense of asymptotic return.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.