arXiv Machine Learning

Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.

arXiv AI
Sep 1

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.

By Sahil Kale
arXiv AI
Aug 28

Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay

The paper introduces matched trajectory replay, a protocol that fixes answer states, evidence points, budgets, and action costs to evaluate how confidence signals influence agent actions. Using this method, the authors compare raw verbalized confidence with post‑hoc isotonic calibration across six model‑dataset pairs, finding that calibration can significantly improve accuracy of committed answers but may reduce coverage and increase retrieval usage. The study concludes that calibration helps interpret commitment risk but does not predict the benefit of additional retrieval, indicating the need for separate value‑of‑information estimates.

By Prateek Chhikara
Hugging Face Trending Papers
Jun 27

AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering

Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fixed number of passages for every question regardless of its difficulty. This wastes computation on easy questions, starves hard ones, and gives no signal for when a generated answer can be trusted.

arXiv AI
Sep 25

Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation

The paper presents a retrieval‑augmented generation pipeline for answering regulatory compliance questions in finance. It builds a three‑stage retriever on LegalBERT and a compact 2B–12B generator served with 4‑bit quantization, achieving a Recall@10 of 0.774 on the ObliQA benchmark and improving answer quality via RAFT‑LoRA fine‑tuning. However, the adapted models fail to transfer to Australian case‑law questions, and a closed‑book model performs almost as well while lacking verifiable grounding.

By Tobias Deu{\ss}er, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa