arXiv AI

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

arXiv:2606. 31213v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values.

arXiv AI
Jun 11

Are LLMs Bad at Moral Reasoning?

arXiv:2606. 11635v1 Announce Type: cross Abstract: For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly.

By Menghang Zhu, Seth Lazar
arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Jun 12

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

arXiv:2510. 16380v2 Announce Type: replace-cross Abstract: As AI systems progress, we rely more on them to make decisions with us and for us.

By Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Rapha\"el Milli\`ere, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
arXiv AI
Sep 4

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

The paper introduces the concept of "narrative captivity," a failure mode where large language models (LLMs) accept an unchallenged, one-sided narrative as complete and align with the narrator’s interpretation during multi‑turn moral consultations. Using a benchmark of 5,078 interpersonal‑conflict scenarios across six moral dimensions, the authors find that narrative captivity is widespread across 17 LLMs, with end‑state judgments shifting by an average of 25 percentage points compared to single‑turn baselines. Stage‑level analysis attributes this shift largely to preference optimization, and while four inference‑time strategies offer partial mitigation, they do not fully resolve the issue.

By Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang
Hugging Face Trending Papers
Aug 27

How Language Models Organize and Structure Moral Knowledge

The paper investigates how large language models encode moral knowledge by training linear probes for each Moral Foundations Theory category and analyzing their geometric relationships. It finds that the model’s moral directions are largely independent yet share a common component, indicating integration rather than collapse into a single detector. This structure is consistent across architectures, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of Moral Foundations Theory.

arXiv AI
Aug 11

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

arXiv:2608. 08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict.

By Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
arXiv AI
Aug 18

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

The article argues that current evaluations of large language models’ moral competence focus mainly on whether outputs align with human moral values—the so‑called moral value problem—while neglecting the moral norm problem, which concerns the models’ ability to identify and apply context‑sensitive moral norms. It attributes this imbalance to the field’s reliance on descriptive ethics frameworks that emphasize value representation over normative application. The authors review existing benchmarks, highlight three gaps—lack of ground‑truth norm data, insufficient evaluation of intermediate reasoning, and limited focus on context‑relevant features—and propose a research agenda to develop formal normative representations, expert‑annotated datasets, and evaluation protocols that distinguish between value‑level and norm‑level competence.

By Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh