arXiv Computation and Language By Ratna Kandala, Niels Vanhasbroeck, Katie Hoemann

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

Read the original on arXiv Computation and Language →

FLAME (FLemish Accounts of Momentary Experiences) is a corpus of nearly 25,000 personal narratives in Belgian-Dutch (Flemish) gathered via experience sampling. The dataset focuses on everyday, culturally grounded themes but presents challenges due to its informal register and low-resource status. Comparative analysis of K‑Means, LDA, and BERTopic shows that BERTopic yields the most coherent and culturally resonant topics according to human evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 3

How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

The paper investigates how Large Language Models (LLMs) construct fictional worlds, specifically examining setting as a measurable aspect of storyworld creation. By generating 1,000 AI stories per model in English and German and comparing them to human-authored fiction from Project Gutenberg, the authors classify narrative space into five categories—action, perceived, visual, descriptive, and no space—using fine‑tuned BERT classifiers. Results show that human texts mainly use action space, grounding narratives in character-environment interaction, while LLMs consistently overproduce perceived space, focusing on atmosphere and affect, with this pattern varying by model and language.

By Katrin Rohrbacher, Bj\"orn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
arXiv AI
Sep 28

Evaluating Cultural Awareness of LLMs for Haitian Creole

The paper presents the first systematic evaluation of cultural awareness in large language models (LLMs) for Haitian Creole, a low‑resource language. Using a benchmark of culturally salient prompts curated by native speakers, the study assesses four dimensions—specificity, bias, diversity, and variation—in a text‑infilling setting. Results show a clear gap between Haitian Creole and higher‑resource French, with Haitian performance more uneven and more affected by French linguistic interference; story generation reveals stereotypical portrayals of Haitian characters even in positive contexts.

By Christelle Clervilsson, Yanzhu Guo
arXiv Computation and Language
Sep 11

Characterizing Narrative Content in Web-scale LLM Pretraining Data

The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.

By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
arXiv Computation and Language
Aug 28

Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.

By Paarth Singh Rathore