arXiv:2606. 04646v1 Announce Type: cross Abstract: Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text.
By Mengao Zhang, Xiang Yang, Chang Liu, Tianhui Tan, Ke-wei Huang
arXiv:2605. 17301v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems implicitly assume mutual consistency among retrieved documents -- an assumption that frequently fails in practice.
By Chenyu Wang, Yueyuan Li, Yingmin Liu, Yang Shu
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surfac...
arXiv:2606. 02093v1 Announce Type: cross Abstract: The task of Error Prediction, namely predicting whether a model output is correct, is commonly tackled with Uncertainty Quantification (UQ).
By Ieva Raminta Stali\=unait\.e, James Bishop, Andreas Vlachos
arXiv:2606. 06197v1 Announce Type: cross Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs).
By Hafez Abdelghaffar, Ahmed Alansary, Ali Hamdi
The paper introduces a semantic correctness taxonomy that categorizes open‑ended QA answers into eight ordered classes, distinguishing between correct, verbose, and hallucinated responses. It releases two datasets—CAP‑Correctness and CAP‑Statements—to support benchmark evaluation and NLI‑based training. The authors also propose CAP (Context‑Aware Precision), a reference‑based metric that scores question‑conditioned statements via bidirectional NLI and demonstrates superior performance under a monotonicity protocol.
By Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
arXiv:2606. 28076v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts.
By Yongxue Shan, Meihan Wu, Cundi Fang, Jie Peng, Xiaodong Wang
arXiv:2607. 16868v1 Announce Type: new Abstract: Large Language Models (LLMs) often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications.
By Yanni Dong, Minghua Liu, Meiling Zhu, Xiaowei Huang, Lijun Zhang
arXiv:2602. 14065v2 Announce Type: replace Abstract: Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval.
By Kai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao, Ye Mo, Liangliang Liu, Haikuan Huang, Bin Li, Jiajun Bu
arXiv:2607. 17742v1 Announce Type: new Abstract: Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability.
By Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, Junbo Zhao
The paper introduces DRAG, a query‑adaptive framework that jointly selects retriever and generator configurations for Retrieval‑Augmented Generation (RAG) systems. Two variants are presented: DRAG_QPP, a training‑free routing method using Query Performance Prediction and perplexity signals, and DRAG_SFT, a supervised approach that fine‑tunes an LLM to predict configurations. Experiments on three LLM families and four QA benchmarks show that DRAG_QPP matches strong static baselines while cutting inference latency, and DRAG_SFT consistently outperforms both static and training‑free adaptive baselines, demonstrating a better effectiveness‑efficiency trade‑off.
By Neeraj Anand, Payel Santra, Partha Basuchowdhuri, Debasis Ganguly, Sumit Bhatia
The paper challenges the common practice of estimating aleatoric uncertainty in large language models (LLMs) by generating multiple clarified inputs and comparing the resulting answers. It argues that answers are unnecessary, costly, and can introduce epistemic leakage, proposing instead a clarification-only method that directly assesses ambiguity from the space of plausible interpretations. Experiments on three benchmarks show the new approach improves AUROC, reduces computational cost, and yields uncertainty estimates less correlated with epistemic uncertainty.
By Omer Nahum, Niv Nayman, Jonathan Fhima, Alon Zolfi, Jeremy Levy, Shai Mazor, Paolo Favaro