arXiv AI By Orion Reblitz-Richardson

How Language Models Organize and Structure Moral Knowledge

Read the original on arXiv AI →

The study investigates how large language models encode moral knowledge by training linear probes for each category of Moral Foundations Theory. It finds that the model’s representations for different moral foundations occupy distinct, largely independent dimensions yet share a common positive component, indicating an integrated but nuanced moral structure. This geometry is consistent across model architectures and scales, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of the theory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 6

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option.