Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608. 12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).
arXiv:2608. 12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).
The study investigates how large language models encode moral knowledge by training linear probes for each category of Moral Foundations Theory. It finds that the model’s representations for different moral foundations occupy distinct, largely independent dimensions yet share a common positive component, indicating an integrated but nuanced moral structure. This geometry is consistent across model architectures and scales, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of the theory.
The paper investigates how large language models encode moral knowledge by training linear probes for each Moral Foundations Theory category and analyzing their geometric relationships. It finds that the model’s moral directions are largely independent yet share a common component, indicating integration rather than collapse into a single detector. This structure is consistent across architectures, emerges early in pre‑training, and reflects corpus statistics rather than the individualizing/binding distinction of Moral Foundations Theory.
arXiv:2609.03330v1 Announce Type: new Abstract: Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems oft...
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
arXiv:2608.29118v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM...
arXiv:2603. 00048v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making.
arXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably lea...
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...
arXiv:2605. 22660v2 Announce Type: replace-cross Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages.
arXiv:2606. 19527v1 Announce Type: new Abstract: Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics?