arXiv Computation and Language

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

FLAME (FLemish Accounts of Momentary Experiences) is a corpus of nearly 25,000 personal narratives in Belgian-Dutch (Flemish) gathered via experience sampling. The dataset focuses on everyday, culturally grounded themes but presents challenges due to its informal register and low-resource status. Comparative analysis of K‑Means, LDA, and BERTopic shows that BERTopic yields the most coherent and culturally resonant topics according to human evaluation.

arXiv Computation and Language
Sep 3

How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

The paper investigates how Large Language Models (LLMs) construct fictional worlds, specifically examining setting as a measurable aspect of storyworld creation. By generating 1,000 AI stories per model in English and German and comparing them to human-authored fiction from Project Gutenberg, the authors classify narrative space into five categories—action, perceived, visual, descriptive, and no space—using fine‑tuned BERT classifiers. Results show that human texts mainly use action space, grounding narratives in character-environment interaction, while LLMs consistently overproduce perceived space, focusing on atmosphere and affect, with this pattern varying by model and language.

By Katrin Rohrbacher, Bj\"orn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
arXiv AI
Sep 28

Evaluating Cultural Awareness of LLMs for Haitian Creole

The paper presents the first systematic evaluation of cultural awareness in large language models (LLMs) for Haitian Creole, a low‑resource language. Using a benchmark of culturally salient prompts curated by native speakers, the study assesses four dimensions—specificity, bias, diversity, and variation—in a text‑infilling setting. Results show a clear gap between Haitian Creole and higher‑resource French, with Haitian performance more uneven and more affected by French linguistic interference; story generation reveals stereotypical portrayals of Haitian characters even in positive contexts.

By Christelle Clervilsson, Yanzhu Guo
arXiv Computation and Language
Sep 11

Characterizing Narrative Content in Web-scale LLM Pretraining Data

The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.

By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
arXiv Computation and Language
Aug 28

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.

By Paarth Singh Rathore
arXiv Computation and Language
Sep 11

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

CHRONOBERG is a temporally structured corpus of English book texts covering 250 years, curated from Project Gutenberg and enriched with temporal annotations. It enables quantification of lexical semantic change via time‑sensitive Valence‑Arousal‑Dominance analysis and the creation of historically calibrated affective lexicons. Experiments show that language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts, highlighting the need for temporally aware training and evaluation pipelines.

By Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski
arXiv AI
Sep 15

SDUs DAISY: A Benchmark for Danish Culture

SDUs DAISY is a factual knowledge benchmark focused on Danish cultural heritage, drawing topics from the Danish Culture Canon 2006. The dataset contains 741 manually verified closed‑ended question‑answer pairs generated by querying Wikipedia pages for each canon artifact, covering a wide historical span from 1300 BCE to contemporary pop music, design, and architecture. Baseline tests with state‑of‑the‑art language models show very low performance (best BLEU 0.17, F1 0.27), highlighting the benchmark’s difficulty and the need for improved cultural knowledge in AI systems.

By Jacob Nielsen, Stine L. Beltoft, Peter Schneider-Kamp, Lukas Galke Poech
arXiv AI
Jun 2

Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

arXiv:2606. 02147v1 Announce Type: cross Abstract: Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation.

By Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina, Ashwath Rao B, Parameswari Krishnamurthy, Muhammad Cendekia Airlangga, Rifo Ahmad Genadi, Nguyen Phan Gia Bao, Amir Hossein Yari, Hawau Olamide Toyin, Nurdaulet Mukhituly, Mena Attia, Besher Hassan, Ahmad Fathan Hidayatullah, Tatsuki Kuribayashi, Haonan Li, Suma Bhat, Fajri Koto
arXiv Computation and Language
Sep 11

From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

The paper introduces a three-tier evaluation framework—recovery, mining, and discovery—for unsupervised narrative label generation in disinformation datasets. It compares clustering-based and graph-community-based pipelines across seven datasets, finding that clustering can underrepresent prominent topics while graph methods produce many singletons that human annotators recognize as valid narratives. The authors release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to aid taxonomy development and dataset expansion.

By Max Upravitelev, Veronika Solopova, Jing Yang, Charlott Jakob, Alexandra Tsiakalou, Neda Foroutan, Vera Schmitt
arXiv Computation and Language
Sep 14

The House with a Million Windows: Interactive Fiction for Narrative Restorying

The House with a Million Windows (HWAMW) is an LLM-based interactive fiction system that lets users narrate a story and then view it through a series of AI-generated "windows" that reframe the narrative in various literary styles. The system is grounded in the psychological restorying intervention, aiming to deepen users' exploration of meaning in their personal stories. Empirical results indicate that HWAMW enhances users' sense of narrative identity, and expert reviews suggest it achieves this by facilitating restorying rather than simply generating new content.

By Cody Kommers, Sarah G Immel, Drew Hemment, Mina Lee