Towards Data Science

From Words to Vectors: What Happens in Between?

The article "From Words to Vectors: What Happens in Between?" explores the process of converting textual data into numerical representations, focusing on techniques such as TF-IDF and vector space models. It discusses how these representations enable text classification tasks and provides a practical overview of the underlying concepts. The piece serves as a guide for readers interested in the mechanics of text preprocessing and feature extraction for machine learning.

Towards Data Science
Aug 20

Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One

The article explains that enterprise document intelligence can be categorized into three distinct corpus types, each requiring a specific architecture. It outlines how to determine the shape of a document collection through three key questions. The piece also discusses the costs associated with building a system for the incorrect corpus type.

By angela shi
Towards Data Science
4d ago

AI Made Data Scientists Faster. Now It’s Expanding the Job.

AI has accelerated data scientists’ productivity, but its influence extends beyond speed. The technology is reshaping who owns data, how judgment is exercised, and the overall career trajectory of data scientists. These changes signal a broader transformation in the field’s structure and responsibilities.

By Yu Dong
Towards Data Science
Aug 26

How Does a RAG Reranker Really Work?

The article "How Does a RAG Reranker Really Work?" explores the inner workings of Retrieval-Augmented Generation (RAG) rerankers, focusing on how data scientists explain the model’s operations behind the scenes. It discusses the impact of these insights on architecture decisions within enterprise document intelligence, specifically in the context of Enterprise Document Intelligence Vol.1 #2D. The piece highlights the importance of transparent model explanations for effective enterprise RAG implementation.

By Kezhan Shi
Towards Data Science
Aug 29

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.

By Kezhan Shi