arXiv AI By Robert Graham, Edward Stevinson, Yariv Barsheshat

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

Read the original on arXiv AI →

arXiv:2607. 14888v1 Announce Type: cross Abstract: Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv AI
Sep 2

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.

By Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein