When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines.
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
By Thennal DK, Hans Ole Hatzel
The paper investigates how Large Language Models (LLMs) construct fictional worlds, specifically examining setting as a measurable aspect of storyworld creation. By generating 1,000 AI stories per model in English and German and comparing them to human-authored fiction from Project Gutenberg, the authors classify narrative space into five categories—action, perceived, visual, descriptive, and no space—using fine‑tuned BERT classifiers. Results show that human texts mainly use action space, grounding narratives in character-environment interaction, while LLMs consistently overproduce perceived space, focusing on atmosphere and affect, with this pattern varying by model and language.
By Katrin Rohrbacher, Bj\"orn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
The House with a Million Windows (HWAMW) is an LLM-based interactive fiction system that lets users narrate a story and then view it through a series of AI-generated "windows" that reframe the narrative in various literary styles. The system is grounded in the psychological restorying intervention, aiming to deepen users' exploration of meaning in their personal stories. Empirical results indicate that HWAMW enhances users' sense of narrative identity, and expert reviews suggest it achieves this by facilitating restorying rather than simply generating new content.
By Cody Kommers, Sarah G Immel, Drew Hemment, Mina Lee
arXiv:2606. 22748v2 Announce Type: replace-cross Abstract: Some professional authors are beginning to use AI tools to help produce their fiction writing.
By Neel Gupta, Maria Antoniak, Melanie Walsh
arXiv:2602. 15851v2 Announce Type: replace-cross Abstract: Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understanding tasks.
By David Y. Liu, Aditya Joshi, Paul Dawson
arXiv:2607. 20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed systems.
By Adam Rigby, Raz Saremi, Azadeh Sohrabinejad, Mehdi Rahimi
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv:2609.14677v1 Announce Type: new
Abstract: Large language models (LLMs) have changed the way people engage with stories. Drawing on public chatbot logs, we can see that when users generate stori...
By Advait Deshmukh, Nora Benedict, Melanie Walsh, Maria Antoniak
arXiv:2609.15998v1 Announce Type: new
Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property...
By Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds
The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.
By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, e...