arXiv AI By Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

Read the original on arXiv AI →

arXiv:2608. 12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
Hugging Face Trending Papers
Aug 27

How Language Models Organize and Structure Moral Knowledge

The paper investigates how large language models encode moral knowledge by training linear probes for each Moral Foundations Theory category and analyzing their geometric relationships. It finds that the model’s moral directions are largely independent yet share a common component, indicating integration rather than collapse into a single detector. This structure is consistent across architectures, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of Moral Foundations Theory.

arXiv AI
Aug 18

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

The article argues that current evaluations of large language models’ moral competence focus mainly on whether outputs align with human moral values—the so‑called moral value problem—while neglecting the moral norm problem, which concerns the models’ ability to identify and apply context‑sensitive moral norms. It attributes this imbalance to the field’s reliance on descriptive ethics frameworks that emphasize value representation over normative application. The authors review existing benchmarks, highlight three gaps—lack of ground‑truth norm data, insufficient evaluation of intermediate reasoning, and limited focus on context‑relevant features—and propose a research agenda to develop formal normative representations, expert‑annotated datasets, and evaluation protocols that distinguish between value‑level and norm‑level competence.

By Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh
arXiv AI
Aug 28

How Language Models Organize and Structure Moral Knowledge

The study investigates how large language models encode moral knowledge by training linear probes for each category of Moral Foundations Theory. It finds that the model’s representations for different moral foundations occupy distinct, largely independent dimensions yet share a common positive component, indicating an integrated but nuanced moral structure. This geometry is consistent across model architectures and scales, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of the theory.

By Orion Reblitz-Richardson