SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 15851v2 Announce Type: replace-cross Abstract: Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understanding tasks.
The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.
The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
arXiv:2608.30754v1 Announce Type: new Abstract: Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine h...
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in...