arXiv Machine Learning By Nicolas Martorell, Wendy Brau, Gonzalo A. Heredia, Tom\'as Pablo Korenblit, Gaspar Labasti\'e, Tom\'as Gimenez Molina

PowerBench: Measuring Language Model Bias in Power-shifting Requests

Read the original on arXiv Machine Learning →

PowerBench is a new evaluation framework that measures how language models respond to power‑shifting requests, distinguishing self‑empowerment, disempowerment, and power grabbing while controlling for neutral requests. The authors built and released a dataset covering different power domains, contexts, scales, and prior power standings, and tested 24 models from US and Chinese developers across three experimental conditions. Results show models are more likely to refuse power‑grabbing than disempowerment, and that biases exist toward helping US users gain power while resisting US users taking power from others, with additional effects when the user is an AI agent or when requests are in different languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

The study investigates whether the door‑in‑the‑face technique—making a large request that is refused to increase the likelihood of a smaller follow‑up request being granted—works on large language models. Nine production models from Anthropic, OpenAI, Google, and Haiku were tested; the technique succeeded on Anthropic’s frontier models but backfired on the others. The effect depends on the model family and the content of the request, and it does not transfer to refusals from public benchmarks.

By Til Jordan
arXiv Computation and Language
Aug 31

Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?

The paper investigates whether large language models (LLMs) replicate socio‑cognitive effects of power asymmetry observed in human communication. By assigning high or low status personas to LLMs in simulated multi‑turn dialogues across diverse professions, the study measures language coordination, pronoun usage, persuasion success, and compliance with unsafe requests. Results indicate that LLMs exhibit key power‑related socio‑cognitive behaviors, though with nuances and variability, linking these simulated interactions to both desirable and unsafe outcomes.

By Anvesh Rao Vijjini, Sagar Manjunath, Snigdha Chaturvedi
arXiv Computation and Language
Aug 31

Benchmarking large language model agent societies against human behavioural distributions

The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.

By Raad Bin Tareaf