arXiv AI

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

arXiv:2609. 11291v1 Announce Type: new Abstract: We post-train Qwen3.

arXiv AI
2d ago

AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content

AfriSyCo investigates how different framing and verification strategies affect the accuracy of language models on African‑language factual content. The study uses a cross‑language factorial design with native‑language follow‑ups and English framing, analyzing 1,415 observations from 100 source questions across seven checkpoints and six languages. Results show that assertive framing boosts target selection by up to 30.4 points, while verification reduces it by 17.4 points, with strong interactions and large variability depending on wording and checkpoint.

By David Ababio Awuni, Rose-Mary Owusuaa Mensah Gyening, Elvis Gyasi Owusu
arXiv Computation and Language
Aug 25

What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

The paper investigates what aspects of language model behavior are controlled by activation steering. By introducing Cross‑Encoding Steering Evaluation, the authors show that steering effects often follow the extraction index of answer identifiers rather than the semantic content of the answers, especially at deeper layers. They also find that a low‑rank output‑sensitive component captures most of this effect, and that different datasets (NormBank, MNLI, SC101) exhibit varying preferences for extraction‑index versus semantic‑label following.

By Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
arXiv AI
Aug 20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.

By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
arXiv Computation and Language
1d ago

Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.

By Levent Bulut