arXiv:2607. 01852v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters.
By Valentin J. J. Kreileder, Johannes Reisinger, Andreas Fischer
Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters. We evaluate if cluster-based semantic chunking improves retrieval and answer quality compared to fixed-size and recursive chunking evaluating on long, structured academic theses using the Retrieval Augmented Generation Assessment (RAGAs) framework.
arXiv:2503. 10677v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language understanding and generation by combining large-scale retrieval systems with generative models.
By Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, Daoyu Wang, Enhong Chen
The paper introduces Knowledge-Aware Semantic Bridging (KASB), a framework designed to improve Retrieval-Augmented Generation (RAG) by aligning the semantic spaces of queries and retrieved contexts. KASB achieves this through intelligent fusion of generative and retrieval-based knowledge in a multistage process, aiming to enhance passage selection quality, relevance, and accuracy. The authors evaluate the method on three popular open-domain Question Answering datasets, demonstrating its effectiveness.
By Xinkai Du, Chao Lv, Yalin Sun, Quanjie Han, Lei Yao, Maosong Sun
The paper introduces INTRA, an attention-based encoder-decoder framework that retrieves directly from its own internal representations instead of using an external retriever. By having decoder attention query pre-encoded evidence chunks, INTRA unifies retrieval and generation, eliminating the typical mismatch seen in retrieval-augmented generation pipelines. Experiments on question-answering benchmarks show that INTRA outperforms strong engineered retrieval pipelines in both evidence recall and overall answer quality.
By Elad Hoffer, Yochai Blau, Edan Kinderman, Ron Banner, Daniel Soudry, Boris Ginsburg
arXiv:2608.21252v1 Announce Type: cross
Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationshi...
By Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han