arXiv AI

Are LLMs Bad at Moral Reasoning?

arXiv:2606. 11635v1 Announce Type: cross Abstract: For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly.

arXiv AI
Jun 12

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

arXiv:2510. 16380v2 Announce Type: replace-cross Abstract: As AI systems progress, we rely more on them to make decisions with us and for us.

By Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Rapha\"el Milli\`ere, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Aug 18

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

The article argues that current evaluations of large language models’ moral competence focus mainly on whether outputs align with human moral values—the so‑called moral value problem—while neglecting the moral norm problem, which concerns the models’ ability to identify and apply context‑sensitive moral norms. It attributes this imbalance to the field’s reliance on descriptive ethics frameworks that emphasize value representation over normative application. The authors review existing benchmarks, highlight three gaps—lack of ground‑truth norm data, insufficient evaluation of intermediate reasoning, and limited focus on context‑relevant features—and propose a research agenda to develop formal normative representations, expert‑annotated datasets, and evaluation protocols that distinguish between value‑level and norm‑level competence.

By Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh
arXiv AI
Sep 7

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.

By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
Hugging Face Trending Papers
Aug 27

How Language Models Organize and Structure Moral Knowledge

The paper investigates how large language models encode moral knowledge by training linear probes for each Moral Foundations Theory category and analyzing their geometric relationships. It finds that the model’s moral directions are largely independent yet share a common component, indicating integration rather than collapse into a single detector. This structure is consistent across architectures, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of Moral Foundations Theory.

arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv AI
Aug 28

How Language Models Organize and Structure Moral Knowledge

The study investigates how large language models encode moral knowledge by training linear probes for each category of Moral Foundations Theory. It finds that the model’s representations for different moral foundations occupy distinct, largely independent dimensions yet share a common positive component, indicating an integrated but nuanced moral structure. This geometry is consistent across model architectures and scales, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of the theory.

By Orion Reblitz-Richardson