arXiv AI By Viktor Vesel\'y, Aleksandar Todorov, Erwan Escudie, Matthia Sabatelli

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

The paper presents a framework that uses saliency maps to create hierarchical attention profiles, tracking how deep reinforcement learning agents allocate attention over time. By comparing these attention trajectories across different conditions and linking them to behavioral metrics, the study reveals algorithm‑specific biases, unintended reward‑driven strategies, and overfitting to redundant sensory inputs. Experiments on Atari 2600 games, custom Pong environments, and biomechanical visuomotor simulations demonstrate that these attention patterns correspond to measurable behavioral differences, establishing attention trajectories as a diagnostic tool beyond traditional performance metrics.

By Charlotte Beylier, Hannah Selder, Arthur Fleig, Simon M. Hofmann, Nico Scherf