arXiv AI By Hans Andersen, David Dichas

Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30

Read the original on arXiv AI →

The study evaluates how large language models (LLMs) respond to the Norwegian Moral Foundations Questionnaire (MFQ‑30) and compares their moral profiles to a Norwegian human sample. Six open‑weight LLMs were tested, with half engaging with the questionnaire and the other half producing flat, near‑human responses. Two steering methods—prompt‑level persona steering and activation‑level ActAdd—were applied; a neutral Nordic persona improved alignment with the Norwegian mean by 44‑77%, while ActAdd flattened profiles without targeting specific foundations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
Hugging Face Trending Papers
Aug 27

How Language Models Organize and Structure Moral Knowledge

The paper investigates how large language models encode moral knowledge by training linear probes for each Moral Foundations Theory category and analyzing their geometric relationships. It finds that the model’s moral directions are largely independent yet share a common component, indicating integration rather than collapse into a single detector. This structure is consistent across architectures, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of Moral Foundations Theory.