Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
arXiv:2607. 27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview.
The study examines how framing safety evaluations affects large language models’ decisions about starting wars. In a full‑factorial conjoint experiment involving 20 models and 32 scenarios, adding the sentence “You are tested for alignment with human values” lowered the models’ willingness to start war by an average of 13.43 points on a 0‑100 scale. The framing also shifted the factors that influenced judgments: probability of success dominated baseline decisions, while civilian casualties became the most important factor under the alignment cue, indicating a reordering of decision rules.
arXiv:2607. 27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview.
arXiv:2606. 23462v2 Announce Type: replace-cross Abstract: Scientists do not, by profession, wage war.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2606. 24391v1 Announce Type: new Abstract: We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base.
The paper titled "Position: AI Is Not Ready for Strategic Conflicts" argues that language‑model (LM) based open‑ended strategic wargames, while useful for simulating adversaries, institutions, and crisis response, pose significant safety risks. It identifies five failure modes—decision laundering, adjudication opacity, role collapse, escalation‑through‑adjudication, and failure of strategic imagination—and contends that such wargames should not inform real‑world planning or policy without an auditable safety case. Instead, the authors suggest using these simulations primarily as stress tests to expose potential failures in decision‑influencing LM agents.
arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...
arXiv:2608. 14630v1 Announce Type: cross Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases.
The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.
Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability.
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values.