arXiv:2606. 03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous tool actions.
By Mengyi Deng, Zhiwei Li, Xin Li, Tingyu Zhu, Ying Zhao, Zhijiang Guo, Wei Wang
arXiv:2608. 15949v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue.
By Cedar Site Bai, Duanshun Li, Zhenyu Liao, Sheikh Sarwar, Huiyuan Chen, Yuan Chen, Changhe Yuan, Haiyang Zhang, Qilin Qi
arXiv:2512.04068v3 Announce Type: replace
Abstract: To handle underspecified or ambiguous queries, AI assistants need a policy for managing their uncertainty to determine (a) when to guess the user i...
By Jonathan Berant, Maximillian Chen, Adam Fisch, Reza Aghajani, Fantine Huot, Mirella Lapata, Jacob Eisenstein
arXiv:2511. 10453v4 Announce Type: replace-cross Abstract: Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong.
By Irina Saparina, Mirella Lapata
arXiv:2609.24290v1 Announce Type: new
Abstract: Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confident...
By Yunxiang Li, Xixin Wu, Helen Meng
The paper presents a tri‑agent framework for evaluating large language models’ question‑clarification abilities. It involves a Question Clarifying Agent that identifies ambiguities and asks follow‑up questions, a Respondent Agent that simulates human replies, and an Evaluator Agent that judges the dialogue using metrics such as ambiguity handling, question quality, dialogue efficiency, language appropriateness, and intent alignment. The authors illustrate the approach with synthetic supply‑chain data and discuss validating the evaluator against human judgments.
By Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra
arXiv:2608. 11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete.
By Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou
PRAGMA is a benchmark designed to evaluate personalized guidance in long‑term conversations. It includes curated longitudinal conversation histories, evidence annotations, and guidance scenarios that reflect evolving user contexts and incorrect assumptions. Experiments show that current retrieval, memory, and long‑context models struggle to recover relevant conversational evidence and to use it effectively for personalized guidance.
By Hyojeong Yu, Hyukhun Koh, Minsung Kim, Yunah Jang, Kyomin Jung
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarificat...
The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.
By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
arXiv:2606. 10156v1 Announce Type: cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace.
By Bharath Sivaram Narasimhan, Karthik R Narasimhan