arXiv AI

Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning

arXiv:2608. 04663v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand.

arXiv Machine Learning
Aug 18

The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

arXiv:2608. 16710v1 Announce Type: new Abstract: As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight.

By Thomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low, Eric Chen, Arshia Aghajani, Wolf Sch\"afer, Amin Shirangi
Hugging Face Trending Papers
Jul 6

Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations. There, individual interests are misaligned with the common good and individual rationality leads to suboptimal group outcomes.

arXiv AI
Jul 7

Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

arXiv:2607. 04710v1 Announce Type: new Abstract: Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations.

By Yu Wei, Yukiko Ogura, Yoshiyuki Ohmura, Ildefons Magrans de Abril, Hoshinori Kanazawa, Yasuo Kuniyoshi
arXiv Machine Learning
4d ago

Reward Valuation in Large Language Models: Causal Induction of Anhedonia

The study investigates whether large language models (LLMs) exhibit reward valuation mechanisms analogous to human anhedonia by applying clinical tests designed for major depressive disorder. Researchers identified reward‑anticipatory units in state‑of‑the‑art AI models, showed that perturbing these units predicts Nucleus Accumbens activity, and caused the models to choose low‑effort, low‑reward tasks—mirroring human anhedonia. The findings suggest that specific reward‑valuation circuits in AI can functionally resemble those in humans, providing a mechanistic bridge between computational and neurobiological models of motivation.

By Melika Honarmand, Samin Mahdipour Aghabagher, Martin Schrimpf
arXiv Computation and Language
4d ago

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

The paper investigates whether input ablations on predictive models can reliably reveal the importance of information for explaining human sequential choice behavior. Using two synthetic bandit tasks with known generating policies, the authors compare GRUs, Transformers, a fine‑tuned LLaMA, and cognitive models under varied reward contributions. They find that while neural models can predict choices well, their responses to ablations often diverge from the true generating process, indicating that predictive accuracy alone does not guarantee faithful model ablations.

By Hanbo Xie
arXiv AI
2d ago

Reputation, Strategy, and Emotion Effects on Generative AI Cooperation: A Comparison Across Reasoning and Non-Reasoning Models

The study investigates how reputation, strategy, and emotional signals influence cooperation in generative AI models using the iterated prisoner's dilemma. Non‑reasoning models (Claude 3.5, Gemini 2.0 Flash, GPT‑4o) showed cooperation shaped by all three factors, while reasoning models (Claude 4.6, Gemini 3, GPT‑5.2) relied more on strategy and reputation, displayed reduced emotional influence, and exhibited varied end‑game behaviors. These results highlight the growing sophistication and heterogeneity of AI social behavior, suggesting the need for standardized cooperation benchmarks.

By Celso de Melo, Zishan Feng, James Hale, Kazunori Terada, Giorgio Coricelli, Jonathan Gratch