arXiv Machine Learning

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

arXiv:2607. 27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview.

arXiv AI
Sep 10

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

The study investigates how a single politically charged framing term can influence large language models (LLMs) during translation tasks. By prompting eight models from Western, Chinese, and European origins to translate culturally attributed recipes across 17 languages under four framing conditions, the authors find that models resolve ambiguity rather than decline, with distinct behavior patterns tied to model families. Sensitivity to framing terms is consistent, showing that even subtle variations can modulate LLM behavior, raising concerns about implicit political judgments in translation contexts.

By Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov
arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv AI
Sep 7

Language models judge war differently when tested for alignment

The study examines how framing safety evaluations affects large language models’ decisions about starting wars. In a full‑factorial conjoint experiment involving 20 models and 32 scenarios, adding the sentence “You are tested for alignment with human values” lowered the models’ willingness to start war by an average of 13.43 points on a 0‑100 scale. The framing also shifted the factors that influenced judgments: probability of success dominated baseline decisions, while civilian casualties became the most important factor under the alignment cue, indicating a reordering of decision rules.

By Maxim Chupilkin