arXiv:2609.34240v2 Announce Type: replace-cross
Abstract: Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, y...
By Jinnuo Liu, Junhao Zhu, Weifeng Jiang, Haoming Liu, Hongyi Wen
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings.
arXiv:2606. 08000v1 Announce Type: cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem.
By Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
The paper introduces a semiotics-aware framework for assessing natural language generation, focusing on how well two texts align in terms of contextual meaning and discourse references. It defines two metrics—Semiotic Fidelity and Semiotic Coverage—to quantify how much of one text’s semiotic profile is supported by the other and how much of the other’s profile is recovered. Experiments reveal that coverage is usually lower than fidelity, and that language models align best with human-curated data at low sampling temperatures, with higher temperatures diminishing this alignment.
By Lorenzo Zangari, Davide Picca
arXiv:2608. 01423v1 Announce Type: cross Abstract: Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response.
By Shengwei Xu, Yuxuan Lu, Yifan Wu, Jason Hartline, Grant Schoenebeck
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
By Thennal DK, Hans Ole Hatzel