arXiv Computation and Language

Reference-Based Analysis of Coherence and Diversity in Open-Ended Text Generation

The paper introduces a reference-based framework to analyze coherence and diversity in open-ended text generation. It evaluates these properties by aligning them with human trajectories, comparing them to human continuations, and estimating their likelihood under a human reference distribution. Experiments show that diversity alignment and mean-based comparisons correlate with human quality ratings, while reference likelihood also associates positively, though results vary by configuration.

arXiv AI
Jun 9

Summarization is Not Dead Yet

arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.

By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
arXiv Computation and Language
Sep 23

A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation

The paper introduces a semiotics-aware framework for assessing natural language generation, focusing on how well two texts align in terms of contextual meaning and discourse references. It defines two metrics—Semiotic Fidelity and Semiotic Coverage—to quantify how much of one text’s semiotic profile is supported by the other and how much of the other’s profile is recovered. Experiments reveal that coverage is usually lower than fidelity, and that language models align best with human-curated data at low sampling temperatures, with higher temperatures diminishing this alignment.

By Lorenzo Zangari, Davide Picca
arXiv Computation and Language
Sep 22

To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews

arXiv:2609.22805v1 Announce Type: new Abstract: Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dia...

By Maitreya Prafulla Chitale, Ketaki Mangesh Shetye, Yash More, Harshit Gupta, Manav Chaudhary, Manish Shrivastava, Vasudeva Varma
arXiv Computation and Language
Sep 1

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

The paper introduces a register-aware framework to evaluate how human-like large language models (LLMs) are, focusing on linguistic feature distributions rather than factual correctness. It uses Maximum Mean Discrepancy (MMD) and 67 Biber lexico‑grammatical features to compare LLM‑generated texts with human reference corpora across different registers. Experiments on seven instruction‑tuned, open‑source models across five English datasets show that all LLMs deviate from human baselines, with closeness to human language varying by register and not by model size.

By Bj\"orn Nieth, Marianna Gracheva, Michaela Mahlberg, Bjoern Eskofier, Emmanuelle Salin
arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
arXiv AI
Aug 26

The Limits of Automatic Evaluation of Creativity in Large Language Models

The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.

By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi