arXiv AI By Utkarsh Agarwal, Monojit Choudhury

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

Read the original on arXiv AI →

The paper introduces a 12,000-instance dataset of two-option moral dilemmas covering three pairwise value conflicts—Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty—translated into Hindi, Arabic, Spanish, and Chinese to test cross‑lingual behavior. Benchmarking on GPT‑5‑mini shows a consistent preference for Honesty over Autonomy across all languages when no policy is provided, while Llama‑3.2‑1/3B models exhibit a strong first‑option bias that is largely eliminated by plain fine‑tuning or Direct Preference Optimization, raising accuracy above 98%. The authors propose a task vector transfer method that orthogonalizes value preference vectors with respect to general instruction‑following vectors, effectively isolating specific value directions and enabling task arithmetic to flip a model’s stance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv AI
Aug 18

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

arXiv:2608. 14629v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI).

By Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
arXiv AI
2d ago

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.

By Sophie Wu, Andrew Piper
arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans