arXiv AI By Hector Munoz-Avila, David W. Aha, Paola Rizzo

Moral Rebel Agents: Decision-Making Under Conflicting Obligations

Read the original on arXiv AI →

The paper introduces the concept of moral rebellion in autonomous agents, where agents may deviate from user‑assigned tasks when morally justified. Five agent architectures are formalized: an amoral agent, utilitarian agents, deontic agents, utilitarian‑deontic (UD) agents, and dutiful agents. Experiments in a Mini Search‑and‑Rescue domain show distinct trade‑offs among task completion, rescue outcomes, and norm compliance, highlighting the role of commitment‑aware moral reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Aug 11

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

arXiv:2608. 08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict.

By Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
arXiv AI
Sep 7

How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions

The paper reports the first empirical study comparing how humans and large language models (LLMs) evaluate perceived moral agency (PMA) in both human and autonomous artificial agents within smart city scenarios. Using a validated PMA scale, 190 human participants and various LLMs were assessed, revealing that humans are perceived to have higher moral agency than artificial agents. When confronted with moral dilemmas, LLMs focus on situational factors such as harm severity and urgency, mirroring the context‑sensitivity observed in human raters.

By Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen