arXiv AI By Thennal DK, Hans Ole Hatzel

Do Large Language Models Always Tell The Same Stories?

Read the original on arXiv AI →

arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

The Limits of Automatic Evaluation of Creativity in Large Language Models

The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.

By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv AI
Sep 21

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.

By Sophie Wu, Andrew Piper
arXiv AI
Jun 4

POLARIS: Guiding Small Models to Write Long Stories

arXiv:2606. 04095v1 Announce Type: cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models.

By Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, John Frederick Wieting
Hugging Face Trending Papers
Aug 6

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines.

arXiv Computation and Language
Sep 3

How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

The paper investigates how Large Language Models (LLMs) construct fictional worlds, specifically examining setting as a measurable aspect of storyworld creation. By generating 1,000 AI stories per model in English and German and comparing them to human-authored fiction from Project Gutenberg, the authors classify narrative space into five categories—action, perceived, visual, descriptive, and no space—using fine‑tuned BERT classifiers. Results show that human texts mainly use action space, grounding narratives in character-environment interaction, while LLMs consistently overproduce perceived space, focusing on atmosphere and affect, with this pattern varying by model and language.

By Katrin Rohrbacher, Bj\"orn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
arXiv AI
Jun 9

Summarization is Not Dead Yet

arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.

By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang