arXiv Machine Learning

Reward Valuation in Large Language Models: Causal Induction of Anhedonia

The study investigates whether large language models (LLMs) exhibit reward valuation mechanisms analogous to human anhedonia by applying clinical tests designed for major depressive disorder. Researchers identified reward‑anticipatory units in state‑of‑the‑art AI models, showed that perturbing these units predicts Nucleus Accumbens activity, and caused the models to choose low‑effort, low‑reward tasks—mirroring human anhedonia. The findings suggest that specific reward‑valuation circuits in AI can functionally resemble those in humans, providing a mechanistic bridge between computational and neurobiological models of motivation.

arXiv Computation and Language
4d ago

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.

By Hanbo Xie
arXiv AI
Jul 7

The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models

arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.

By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang