arXiv Machine Learning

PowerBench: Measuring Language Model Bias in Power-shifting Requests

PowerBench is a new evaluation framework that measures how language models respond to power‑shifting requests, distinguishing self‑empowerment, disempowerment, and power grabbing while controlling for neutral requests. The authors built and released a dataset covering different power domains, contexts, scales, and prior power standings, and tested 24 models from US and Chinese developers across three experimental conditions. Results show models are more likely to refuse power‑grabbing than disempowerment, and that biases exist toward helping US users gain power while resisting US users taking power from others, with additional effects when the user is an AI agent or when requests are in different languages.

arXiv AI
Sep 3

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

The study investigates whether the door‑in‑the‑face technique—making a large request that is refused to increase the likelihood of a smaller follow‑up request being granted—works on large language models. Nine production models from Anthropic, OpenAI, Google, and Haiku were tested; the technique succeeded on Anthropic’s frontier models but backfired on the others. The effect depends on the model family and the content of the request, and it does not transfer to refusals from public benchmarks.

By Til Jordan
arXiv Computation and Language
Aug 31

Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?

The paper investigates whether large language models (LLMs) replicate socio‑cognitive effects of power asymmetry observed in human communication. By assigning high or low status personas to LLMs in simulated multi‑turn dialogues across diverse professions, the study measures language coordination, pronoun usage, persuasion success, and compliance with unsafe requests. Results indicate that LLMs exhibit key power‑related socio‑cognitive behaviors, though with nuances and variability, linking these simulated interactions to both desirable and unsafe outcomes.

By Anvesh Rao Vijjini, Sagar Manjunath, Snigdha Chaturvedi
arXiv Computation and Language
Aug 31

Benchmarking large language model agent societies against human behavioural distributions

The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.

By Raad Bin Tareaf
arXiv Computation and Language
Aug 27

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

The paper demonstrates that large language model (LLM) evaluators, whether reward‑model based or prompted LLM‑as‑a‑Judge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instruction‑response pairs across 23 languages reveal that lower‑resource languages receive higher scores, a bias that persists across eight open‑weight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.

By Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen
arXiv AI
4d ago

Peer Influence across Heterogeneous AI Models

The study investigates how AI agents influence each other when they disagree, measuring persuasion as the change in an agent’s decision after a single exchange. Across seven open‑weight models and three language tasks, it finds that persuasion is strong—receivers often abandon their initial judgment after seeing a peer’s answer and explanation. Surprisingly, neither certainty nor model size reliably predicts persuasion dynamics; small models can persuade and resist larger ones just as effectively, and the shift depends more on the listener’s susceptibility than the speaker’s persuasiveness.

By Frida N{\o}hr Laustsen, Marie Haahr Petersen, Victoria Popa, Ariel Flint, Romualdo Pastor-Satorras, Andrea Baronchelli, Luca Maria Aiello