The House with a Million Windows (HWAMW) is an LLM-based interactive fiction system that lets users narrate a story and then view it through a series of AI-generated "windows" that reframe the narrative in various literary styles. The system is grounded in the psychological restorying intervention, aiming to deepen users' exploration of meaning in their personal stories. Empirical results indicate that HWAMW enhances users' sense of narrative identity, and expert reviews suggest it achieves this by facilitating restorying rather than simply generating new content.
By Cody Kommers, Sarah G Immel, Drew Hemment, Mina Lee
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
By Thennal DK, Hans Ole Hatzel
arXiv:2608. 15654v1 Announce Type: cross Abstract: Large language models can write fluent stories, but open-ended storytelling requires more than local fluency.
By Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Ka Man Yan, Ying Li
When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines.
The paper investigates how characters in large language model (LLM)-generated stories compare to those in human-written stories. Using narratological definitions, it analyzes eight complex character dimensions—including stylization and wholeness—to automatically categorize characters in both LLM and human texts. The study then contrasts these categories to answer whether LLMs produce similar and varied character portrayals as human authors.
By Anneliese Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi
arXiv:2602. 15851v2 Announce Type: replace-cross Abstract: Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understanding tasks.
By David Y. Liu, Aditya Joshi, Paul Dawson
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv:2607. 00009v1 Announce Type: cross Abstract: Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure.
By Mingzhe Lu, Yanbing Liu, Jiayue Wu, Jiarui Zhang, Qihao Wang, Yue Hu, Yunpeng Li, Yangyan Xu
arXiv:2601.15295v2 Announce Type: replace-cross
Abstract: Interactive narrative (IN) authors craft spaces of divergent narrative possibilities for players to explore, with the player's input determin...
By Yi Wang, John Joon Young Chung, Melissa Roemmele, Yuqian Sun, Tiffany Wang, Shm Garanganao Almeda, Brett A. Halperin, Yuwen Lu, Max Kreminski
arXiv:2606. 17391v1 Announce Type: cross Abstract: Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail.
By Logan Mann, Abdur Rahman, Mohammad Saifullah, Taaha Kazi, Vasu Sharma
arXiv:2608. 19437v1 Announce Type: cross Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality.
By Nirav Patel, Josiah Crossman, Eva Aggarwal, Emily Wenger
arXiv:2605. 17064v2 Announce Type: replace Abstract: Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing.
By Jan Zierstek, Matteo Batelic, Maya Medjad, Tim Sch\"onenberger