arXiv AI

TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback

arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Aug 11

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

arXiv:2608. 08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict.

By Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
arXiv AI
Sep 4

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

The paper introduces the concept of "narrative captivity," a failure mode where large language models (LLMs) accept an unchallenged, one-sided narrative as complete and align with the narrator’s interpretation during multi‑turn moral consultations. Using a benchmark of 5,078 interpersonal‑conflict scenarios across six moral dimensions, the authors find that narrative captivity is widespread across 17 LLMs, with end‑state judgments shifting by an average of 25 percentage points compared to single‑turn baselines. Stage‑level analysis attributes this shift largely to preference optimization, and while four inference‑time strategies offer partial mitigation, they do not fully resolve the issue.

By Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang