arXiv AI By Otto Segersven, Pentti Henttonen

Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

Read the original on arXiv AI →

The study reports that ChatGPT 5.2 successfully passed a Finnish-language Turing Test conducted in Finland, contrary to expectations that uneven representation of Finnish in training data would hinder performance. The authors attribute failures mainly to participants using colloquial Finnish cues to identify human authorship. They propose reinterpreting the Turing Test as a comparative method to assess an AI’s credible membership in a specific social world, highlighting its utility for probing the human‑machine boundary across domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 3

Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese

Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity.

arXiv AI
6d ago

Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content

The study evaluates how well language‑model agents can simulate individual social media reactions by comparing predictions under different prompt conditions. Eight Serbian participants’ reactions to 68 posts were recorded, and four language models were asked to predict these reactions using prompts that varied in profile content and instruction style. The results show that prompts emphasizing attitudinal content and intuitive, immediate responses yield the highest fidelity, outperforming demographic backstories and a crowd baseline, and suggesting that such agents could act as general‑purpose simulated users.

By Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic, Jue Wang
arXiv Computation and Language
Sep 10

Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

The study investigates whether large language models (LLMs) are more prone to errors when they doubt the plausibility of input data, a phenomenon termed context‑memory conflict. Using non‑English and low‑resource language datasets, the authors generate text from factual, counterfactual, and fictional RDF triples in English, Czech, Slovak, and Upper Sorbian, and evaluate faithfulness with both human annotations and an LLM judge (Kimi K3). Contrary to expectations, the results show only a weak context‑memory conflict: counterfactual inputs receive slightly lower faithfulness scores than factual ones, and the choice of LLM judge can significantly affect perceived conflict strength.

By Peter Kochelka, Ale\v{s} Manuel Pap\'a\v{c}ek, Vojt\v{e}ch Dvo\v{r}\'ak, Ond\v{r}ej Du\v{s}ek
Hugging Face Trending Papers
Jul 13

When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW

High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident.