arXiv Machine Learning

Calibrated Surprise: An Information-Theoretic Account of Creative Quality

arXiv:2604. 26269v2 Announce Type: replace-cross Abstract: In the era of large language models, creative writing quality lacks a computable theoretical anchor.

arXiv AI
Aug 26

The Limits of Automatic Evaluation of Creativity in Large Language Models

The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.

By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv AI
Jun 4

POLARIS: Guiding Small Models to Write Long Stories

arXiv:2606. 04095v1 Announce Type: cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models.

By Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, John Frederick Wieting
Hugging Face Trending Papers
Aug 3

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias.

arXiv AI
Sep 24

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

arXiv:2609. 28245v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored.

By AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States), Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)
arXiv AI
Aug 20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.

By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
arXiv Computation and Language
Sep 10

AI translation of literary texts is "fine", but readers still prefer human translations

arXiv:2606.26040v2 Announce Type: replace Abstract: AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know enough about how readers expe...

By Yves Ferstler, Adam Podoxin, Ty Brassington, Ga\"elle Laperri\`ere, Roman Grundkiewicz, Marie-Jean Meurs, Maite Taboada, Marzena Karpinska