arXiv AI

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

arXiv:2606. 06197v1 Announce Type: cross Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs).

arXiv AI
6d ago

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

The paper introduces a knowledge‑graph‑based evaluation framework, S3KG, to assess whether large language models truly understand context in question answering tasks. S3KG combines structural and semantic signals into a single similarity score and is paired with a diagnostic analysis that pinpoints reasoning errors at the triplet level. Across nine benchmarks, the method outperforms existing baselines, achieving up to +7.6 F1 points and an AUROC of 0.973.

By Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Kamal Premaratne, Uthayasanker Thayasivam
arXiv AI
Aug 11

KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs

arXiv:2608. 09779v1 Announce Type: cross Abstract: Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform.

By Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu, Fiona Veazey, Peter Hubbert, Hui Su, Paul Buitelaar
arXiv Computation and Language
Aug 24

Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering

Granuscore is a reference‑free metric that measures the granularity of text by exploiting the structure of a hierarchical embedding space. It successfully reproduces known hierarchical orderings on the Granola‑EQ dataset, distinguishes granularity across different discourse contexts, and explains sentence‑specificity variations beyond sentence length. The authors also apply Granuscore to four question‑answering benchmarks, revealing systematic differences in granularity among questions, gold answers, and model outputs, thereby offering a new lens for assessing QA dataset difficulty.

By Lukas Ellinger, Alexander Fichtl, Miriam Ansch\"utz, Georg Groh
arXiv Computation and Language
Aug 31

PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering

PRISM is an agentic retrieval framework that uses large language models in a structured loop to improve evidence gathering for multi‑hop question answering. It splits retrieval into three specialized agents—a Question Analyzer, a Selector focused on precision, and an Adder focused on recall—whose iterative interaction yields a compact yet comprehensive evidence set. Experiments on HotpotQA, 2WikiMultiHopQA, MuSiQue, and MultiHopRAG show that PRISM consistently outperforms strong baselines by achieving higher retrieval accuracy and filtering out distracting content.

By Md Mahadi Hasan Nahid, Davood Rafiei
arXiv AI
Jul 15

CANDI: Contextual Alignment for Niche Domains Question Answering

arXiv:2607. 11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge.

By Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth