Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia
arXiv:2607. 06626v1 Announce Type: new Abstract: Recent Vision-Language Models capture increasingly complex aspects of human cognition.
The study investigates whether large language models (LLMs) exhibit reward valuation mechanisms analogous to human anhedonia by applying clinical tests designed for major depressive disorder. Researchers identified reward‑anticipatory units in state‑of‑the‑art AI models, showed that perturbing these units predicts Nucleus Accumbens activity, and caused the models to choose low‑effort, low‑reward tasks—mirroring human anhedonia. The findings suggest that specific reward‑valuation circuits in AI can functionally resemble those in humans, providing a mechanistic bridge between computational and neurobiological models of motivation.
arXiv:2607. 06626v1 Announce Type: new Abstract: Recent Vision-Language Models capture increasingly complex aspects of human cognition.
The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.
arXiv:2607. 07753v1 Announce Type: cross Abstract: Modelling psychological disorders in artificial agents offers both a testbed for computational psychiatry and a lens on the failure modes of affective control.
arXiv:2608. 05111v1 Announce Type: new Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies.
arXiv:2604. 19139v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI, a growing and increasingly conspicuous phenomenon has emerged: the proliferation of verbal tics--repetitive, formulaic linguistic patterns that pervade model outputs.
arXiv:2606. 15507v1 Announce Type: new Abstract: Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it.
arXiv:2607. 04590v1 Announce Type: new Abstract: Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences.
arXiv:2609.27532v1 Announce Type: new Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The...
arXiv:2609.35599v2 Announce Type: replace Abstract: When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise thei...
arXiv:2608. 04663v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand.
arXiv:2607. 18966v1 Announce Type: new Abstract: Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective.
arXiv:2609.01170v1 Announce Type: new Abstract: Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organizati...