Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces a reference-based framework to analyze coherence and diversity in open-ended text generation. It evaluates these properties by aligning them with human trajectories, comparing them to human continuations, and estimating their likelihood under a human reference distribution. Experiments show that diversity alignment and mean-based comparisons correlate with human quality ratings, while reference likelihood also associates positively, though results vary by configuration.
arXiv:2607. 05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes.
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
arXiv:2606. 08417v1 Announce Type: cross Abstract: Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling.
arXiv:2509. 25359v2 Announce Type: replace-cross Abstract: We present a systematic stress-test of geometric metrics for LLM evaluation.
arXiv:2603.04419v3 Announce Type: replace-cross Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify...