arXiv:2606. 09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is verifiable.
By Liyi Zhang, Akshay K. Jagadish, Brenden M. Lake, Thomas L. Griffiths
arXiv:2606. 08497v1 Announce Type: new Abstract: As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability.
By Minyoung Hwang, Seokhyun Lee, Changhee Lee
arXiv:2606. 20245v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive parametric knowledge and in-context learning ability, enabling them to incorporate external information provided in the input prompt.
By Huang Peng, Jiuyang Tang, Weixin Zeng, Hao Xu, Xiang Zhao
arXiv:2510. 11194v3 Announce Type: replace Abstract: Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions.
By Peiming Li, Zhiyuan Hu, Yang Tang, Shiyu Li, Xi Chen
The paper presents a tri‑agent framework for evaluating large language models’ question‑clarification abilities. It involves a Question Clarifying Agent that identifies ambiguities and asks follow‑up questions, a Respondent Agent that simulates human replies, and an Evaluator Agent that judges the dialogue using metrics such as ambiguity handling, question quality, dialogue efficiency, language appropriateness, and intent alignment. The authors illustrate the approach with synthetic supply‑chain data and discuss validating the evaluator against human judgments.
By Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra
arXiv:2606. 03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous tool actions.
By Mengyi Deng, Zhiwei Li, Xin Li, Tingyu Zhu, Ying Zhao, Zhijiang Guo, Wei Wang
arXiv:2609.26610v1 Announce Type: new
Abstract: Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested...
By David Torres-Moreno, Jorge Hermosillo-Valadez
arXiv:2606. 12332v1 Announce Type: cross Abstract: Evaluating multi-turn dialogue is challenging because quality emerges across turns rather than within individual responses.
By Paul He, Shiva Kasiviswanathan, Dominik Janzing
arXiv:2609.13520v1 Announce Type: new
Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they a...
By Kyle Richardson, Cullen Anderson, Pranav Balakrishnan, Takuto Ban, Daksha Ladia, Ankita Gupta, Marisa Hudspeth
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarificat...
arXiv:2606. 02093v1 Announce Type: cross Abstract: The task of Error Prediction, namely predicting whether a model output is correct, is commonly tackled with Uncertainty Quantification (UQ).
By Ieva Raminta Stali\=unait\.e, James Bishop, Andreas Vlachos
arXiv:2607. 06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English.
By Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico