arXiv AI

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

arXiv:2608. 11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined.

arXiv AI
Sep 10

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

The study investigates how a single politically charged framing term can influence large language models (LLMs) during translation tasks. By prompting eight models from Western, Chinese, and European origins to translate culturally attributed recipes across 17 languages under four framing conditions, the authors find that models resolve ambiguity rather than decline, with distinct behavior patterns tied to model families. Sensitivity to framing terms is consistent, showing that even subtle variations can modulate LLM behavior, raising concerns about implicit political judgments in translation contexts.

By Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov
arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv Computation and Language
Sep 18

Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.

By Levent Bulut
arXiv Computation and Language
Aug 31

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.

By Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka
arXiv Computer Vision
Aug 28

Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

The paper introduces CMPM, a Chinese Multi-Panel Meme benchmark comprising 1,214 annotated samples that capture five structural types, ordering dependencies, panel-order constraints, and optional comment context. It defines a two-layer evaluation: Task 1 tests structure typing and order-sensitive panel sequencing, while Task 2 assesses Chinese meme explanation generation using human ratings across visual, panel, humor, context, and faithfulness dimensions. Benchmarking five large vision‑language models shows that accuracy on canonical displays does not guarantee order understanding, as performance drops sharply under shuffled conditions, and that Gemini 3.1 Pro and GPT‑5.5 outperform open models in Task 2, with comment context providing only modest gains.

By Haihan Li, Haihao Li, Zhenfei Xu, Jize Qian
arXiv AI
Sep 12

Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems

The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.

By Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua, Azza Bouleimen, Robin Forsberg, Corinna Hertweck, Stefania Ionescu, Nicol\`o Pagan, Ancsa Hannak, Joachim Baumann