The paper "Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems" investigates how commercial text‑to‑image models silently modify user prompts before generating images, a step that is often hidden from users. Using the multilingual benchmark WORLDVIEW, the authors audit the revision layer in DALL‑E‑3, Imagen‑4, and GPT‑Image‑1.5, finding that non‑Western and non‑Anglophone contexts are disproportionately marked, reduced to narrow vocabularies, and stereotyped. The study demonstrates that the revision layer itself is a previously undocumented causal source of cultural stereotyping, underscoring the need to audit deployed systems rather than just the underlying models.
By Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua, Azza Bouleimen, Robin Forsberg, Corinna Hertweck, Stefania Ionescu, Nicol\`o Pagan, Ancsa Hannak, Joachim Baumann
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
arXiv:2609.37287v1 Announce Type: cross
Abstract: Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preserva...
By Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, Tao Chen
arXiv:2608.30541v1 Announce Type: new
Abstract: Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent...
By Ran Zhang, Miryam de Lhoneux, Wessel Poelman
arXiv:2604. 18347v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) achieved rapid progress in the recent years.
By Daniela Baiamonte, Elena Fano, Matteo Gabburo, Stefano Simonazzi, Leonardo Rigutini, Andrea Zugarini
The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.
By Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett