arXiv:2607. 29334v1 Announce Type: cross Abstract: Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty.
By Ningzhi Liu, Yannic Hinrichs, Jonas R. Kunst
The study investigates how a single politically charged framing term can influence large language models (LLMs) during translation tasks. By prompting eight models from Western, Chinese, and European origins to translate culturally attributed recipes across 17 languages under four framing conditions, the authors find that models resolve ambiguity rather than decline, with distinct behavior patterns tied to model families. Sensitivity to framing terms is consistent, showing that even subtle variations can modulate LLM behavior, raising concerns about implicit political judgments in translation contexts.
By Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov
arXiv:2508.16013v2 Announce Type: replace
Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...
By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.
By Levent Bulut
arXiv:2604. 14180v2 Announce Type: replace-cross Abstract: We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.
By Jiuting Chen, Yuan Lian, Hao Wu, Tianqi Huang, Hiroshi Sasaki, Makoto Kouno, Jongil Choi
arXiv:2607. 27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview.
By Haran Shani-Narkiss, Michael Fire, Oren Tsur
CNeo-Bench is a new benchmark comprising 4,759 Chinese neologisms, each with reference definitions and categorized by linguistic mechanisms such as phonetic substitution and visual character decomposition. The benchmark includes a two-tier evaluation framework that tests whether models can describe a neologism and whether they can manipulate its underlying mechanism. Evaluation of 18 large language models shows that most perform poorly on definition generation (below 40%) and exhibit a recognition‑manipulation gap, often paraphrasing rather than restoring the original form; few‑shot prompting helps but does not fully resolve the errors.
By Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka
The paper introduces CMPM, a Chinese Multi-Panel Meme benchmark comprising 1,214 annotated samples that capture five structural types, ordering dependencies, panel-order constraints, and optional comment context. It defines a two-layer evaluation: Task 1 tests structure typing and order-sensitive panel sequencing, while Task 2 assesses Chinese meme explanation generation using human ratings across visual, panel, humor, context, and faithfulness dimensions. Benchmarking five large vision‑language models shows that accuracy on canonical displays does not guarantee order understanding, as performance drops sharply under shuffled conditions, and that Gemini 3.1 Pro and GPT‑5.5 outperform open models in Task 2, with comment context providing only modest gains.
By Haihan Li, Haihao Li, Zhenfei Xu, Jize Qian
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
arXiv:2607. 22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior.
By Viktoriia Makovska, George Fletcher
The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.
By Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua, Azza Bouleimen, Robin Forsberg, Corinna Hertweck, Stefania Ionescu, Nicol\`o Pagan, Ancsa Hannak, Joachim Baumann
arXiv:2608. 12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.
By Arnav Srivastav