arXiv Machine Learning

Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia

arXiv:2607. 06626v1 Announce Type: new Abstract: Recent Vision-Language Models capture increasingly complex aspects of human cognition.

arXiv Machine Learning
4d ago

Reward Valuation in Large Language Models: Causal Induction of Anhedonia

The study investigates whether large language models (LLMs) exhibit reward valuation mechanisms analogous to human anhedonia by applying clinical tests designed for major depressive disorder. Researchers identified reward‑anticipatory units in state‑of‑the‑art AI models, showed that perturbing these units predicts Nucleus Accumbens activity, and caused the models to choose low‑effort, low‑reward tasks—mirroring human anhedonia. The findings suggest that specific reward‑valuation circuits in AI can functionally resemble those in humans, providing a mechanistic bridge between computational and neurobiological models of motivation.

By Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf
arXiv Computation and Language
4d ago

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.

By Hanbo Xie
arXiv AI
Jul 24

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.

By Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, Ranjay Krishna
arXiv Machine Learning
Sep 4

Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

The paper presents a framework that uses saliency maps to create hierarchical attention profiles, tracking how deep reinforcement learning agents allocate attention over time. By comparing these attention trajectories across different conditions and linking them to behavioral metrics, the study reveals algorithm‑specific biases, unintended reward‑driven strategies, and overfitting to redundant sensory inputs. Experiments on Atari 2600 games, custom Pong environments, and biomechanical visuomotor simulations demonstrate that these attention patterns correspond to measurable behavioral differences, establishing attention trajectories as a diagnostic tool beyond traditional performance metrics.

By Charlotte Beylier, Hannah Selder, Arthur Fleig, Simon M. Hofmann, Nico Scherf