arXiv AI

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

arXiv:2607. 01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety.

arXiv AI
Sep 10

Attribution in Scientific Literature: New Benchmark and Methods

The paper introduces REASONS, a benchmark of 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution under different evidence conditions. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to balance reliability and responsiveness. Experiments with proprietary and open-source LLMs across various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but increases abstention, while adversarial metadata can push hallucination rates above 85%. Human evaluation confirms a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.

By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur
arXiv Computer Vision
Sep 18

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large‑scale benchmark for fine‑grained, element‑level source attribution in Document Visual Question Answering (VQA). It introduces MAPPET, a Mask‑based Perplexity‑Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question‑answer pairs with element‑level grounding, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality, revealing that even strong models often fail to localize supporting elements.

By Luca De Grandis (University of Modena and Reggio Emilia, Modena, Italy), Silvia Cappelletti (University of Modena and Reggio Emilia, Modena, Italy), William Raccagni (University of Modena and Reggio Emilia, Modena, Italy, University of Pisa, Pisa, Italy), Marcella Cornia (University of Modena and Reggio Emilia, Modena, Italy), Lorenzo Baraldi (University of Modena and Reggio Emilia, Modena, Italy), Rita Cucchiara (University of Modena and Reggio Emilia, Modena, Italy)
arXiv AI
Sep 17

Abstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations

The paper introduces REASONS, a benchmark comprising 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution by large language models. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to assess the trade-off between reliability and responsiveness. Experiments on proprietary and open-source LLMs under various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but may increase abstention, while retrieval-augmented variants often maintain near-zero abstention. Human evaluation reveals a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.

By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur
arXiv AI
Jun 16

VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA

arXiv:2606. 16092v1 Announce Type: cross Abstract: Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements.

By Young Rok Jang, Hyesoo Kong, Kyunghwan An, Jae Sub Huh, Gyeonghun Kim, Stanley Jungkyu Choi
arXiv Computation and Language
Sep 11

Probing for Knowledge Attribution in Large Language Models

The paper introduces a method for identifying the dominant knowledge source behind large language model (LLM) outputs, distinguishing between faithfulness violations (misuse of provided context) and factuality violations (errors in internal knowledge). A simple linear probe trained on hidden representations can reliably classify this source, and the authors present AttriWiki, a self‑supervised pipeline that generates labeled training data by prompting models to recall withheld entities or read them from context. Probes trained on AttriWiki achieve high Macro‑F1 scores across several models and datasets, generalize zero‑shot to a benchmark, and show that attribution mismatches can increase error rates by up to 70%. "whyItMatters":"The study demonstrates that knowing the source of an LLM’s answer is crucial for effective mitigation of hallucinations, as attribution mismatches significantly raise error rates."

By Ivo Brink, Alexander Boer, Dennis Ulmer
arXiv AI
Aug 18

SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling

arXiv:2508. 09105v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) significantly improve the knowledge coverage and contextual understanding of Large Language Models (LLMs) by introducing external knowledge sources.

By Shixuan Sun, Siyuan Liang, Jianjie Huang, Jingzhi Li, Xiaochun Cao
Hugging Face Trending Papers
Sep 17

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large-scale benchmark that provides fine-grained, element-level source attribution for Document Visual Question Answering (VQA). It introduces MAPPET, a Mask-based Perplexity-Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question-answer pairs with grounding annotations, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality.