Towards Data Science

Embeddings Aren’t Magic: The Predictable Failure Modes of RAG Retrieval

Enterprise Document Intelligence [Vol. 1 #2] Why the same vector search that handles synonyms and paraphrase silently fails on negation, exact identifiers, and your company’s acronyms, and what to use when it does.

Towards Data Science
Aug 20

Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One

The article explains that enterprise document intelligence can be categorized into three distinct corpus types, each requiring a specific architecture. It outlines how to determine the shape of a document collection through three key questions. The piece also discusses the costs associated with building a system for the incorrect corpus type.

By angela shi
Towards Data Science
Aug 29

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.

By Kezhan Shi
Towards Data Science
Sep 23

From Words to Vectors: What Happens in Between?

The article "From Words to Vectors: What Happens in Between?" explores the process of converting textual data into numerical representations, focusing on techniques such as TF-IDF and vector space models. It discusses how these representations enable text classification tasks and provides a practical overview of the underlying concepts. The piece serves as a guide for readers interested in the mechanics of text preprocessing and feature extraction for machine learning.

By Nikhil Dasari
arXiv AI
Sep 3

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search describes DocuSearch, an offline multi‑agent system designed for telecom network operations. The system combines semantic vector search, BM25 full‑text search, and knowledge‑graph neighbor expansion, merges the results via Reciprocal Rank Fusion, and reranks with a cross‑encoder before pruning with Maximal Marginal Relevance. A per‑chunk evaluation loop ensures only grounded answers are returned, achieving Precision@10 of 0.69, Recall@10 of 0.79, and an 89.6% grounding rate—improvements of 15, 16, and 18.4 percentage points over a dense‑only baseline.

By Harish Saragadam, Sudhanshu Sharma, Meghana Pujari
arXiv AI
Aug 25

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

arXiv:2608.21702v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity,...

By Jing Liu, Yongxing Qi, Muchen Jiang, Chengnan Hu, Qingqing Peng, Haoming Wang, Yuqing Wang, Yang Yu, Xu Zhang, Ting Wu
Towards Data Science
Aug 30

Noisy Text in RAG: Typos, OCR, and the Gap Classical Spell-Check Leaves

The article discusses how noisy text—stemming from user typos, rapid transcription errors, and OCR character mistakes—poses challenges for Retrieval-Augmented Generation (RAG) systems. It explains that traditional spell-checking only addresses one type of error, while embeddings are needed to handle the remaining noise. The piece highlights the need for more robust solutions in enterprise document intelligence.

By Kezhan Shi