PragAlign is a reply‑assistance system that separates context reading from selective clarification, designed for multilingual settings. In evaluations with native Chinese and Japanese speakers, PragAlign achieved better rankings than baseline methods in Chinese and the highest top‑rank rate in Japanese, while both groups often chose the same top condition. The study highlights both shared and language‑specific judgment patterns that can inform culturally aware reply‑assistance tools.
By Xin Zhong, Satori Hachisuka
The study reports that ChatGPT 5.2 successfully passed a Finnish-language Turing Test conducted in Finland, contrary to expectations that uneven representation of Finnish in training data would hinder performance. The authors attribute failures mainly to participants using colloquial Finnish cues to identify human authorship. They propose reinterpreting the Turing Test as a comparative method to assess an AI’s credible membership in a specific social world, highlighting its utility for probing the human‑machine boundary across domains.
By Otto Segersven, Pentti Henttonen
arXiv:2608. 14896v1 Announce Type: cross Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it.
By Florian Braun
Large language models (LLMs) are increasingly deployed in hiring workflows, yet most research on gender bias in LLM hiring decisions has focused on English-language, Western-format resumes. This study examines whether pro-female gender bias extends to a Japanese corporate context and evaluates two practical mitigation strategies.
arXiv:2606. 01260v1 Announce Type: cross Abstract: Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape.
By Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto
The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.
By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados