arXiv:2608. 05726v1 Announce Type: cross Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts.
By Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn
arXiv:2606. 08417v1 Announce Type: cross Abstract: Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling.
By Antonio Franca, Alexander Tong
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of the evaluated text, which is known as scoring bias.
arXiv:2606. 07537v1 Announce Type: cross Abstract: Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales.
By Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa Naeem
arXiv:2607. 03882v1 Announce Type: cross Abstract: LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language.
By Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
arXiv:2606. 05486v1 Announce Type: cross Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, while existing attribution methods are designed to explain observable outputs such as logits or generated tokens.
By Govind Ramesh, Yao Dou, Wei Xu