arXiv AI By Hortense Fong, George Gui, Bo Yang

Modeling Story Expectations: A Generative Framework using LLMs

Read the original on arXiv AI →

arXiv:2412. 15239v4 Announce Type: replace-cross Abstract: Consumers' engagement with stories is shaped by their expectations about what will happen next, yet modeling these forward-looking beliefs over unstructured narrative content has remained challenging.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

Characterizing Narrative Content in Web-scale LLM Pretraining Data

The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.

By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak
arXiv AI
Aug 19

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

The paper investigates how large language models (LLMs) engage with long-form narratives by comparing their generated novel summaries to human-authored ones. Researchers align sentences from 150 human-written summaries to specific chapters, highlighting the challenge of this alignment task and the complexity of summarization. They find stylistic differences and that LLMs tend to focus more on the ends of texts, suggesting insights into why models may struggle with narrative comprehension.

By Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan