Towards Data Science

Building Trustworthy Production RAG Systems Through Continuous Evaluation

A practical guide to building an evaluation workflow that catches retrieval failures, hallucinations, and performance drift before they reach users The post Building Trustworthy Production RAG Systems Through Continuous Evaluation appeared first on Towards Data Science .

Towards Data Science
Sep 22

Break Your Own RAG Pipeline Before Users Do

The article discusses a small adversarial test set designed to detect retrieval failures in Retrieval-Augmented Generation (RAG) pipelines that typical evaluation sets might miss. It emphasizes the importance of proactively testing your own RAG system to uncover hidden weaknesses before users encounter them. By using this targeted test set, developers can improve the reliability and robustness of their RAG models.

By Sara Nobrega
Towards Data Science
Jun 9

10 Common RAG Mistakes We Keep Seeing in Production

Enterprise Document Intelligence [Vol. 1 #4bis] - A coauthor note on the brick-by-brick pitfalls that justified the four-brick split, before Part II walks the fixes The post 10 Common RAG Mistakes We Keep Seeing in Production appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Aug 31

Why RAG Complexity Should Be Earned

The article outlines a framework for constructing Retrieval-Augmented Generation (RAG) pipelines that progressively add complexity as needed to address observed failure modes. It begins with basic lexical and hybrid search techniques, then incorporates reranking and agentic information‑seeking strategies to improve performance. The approach emphasizes that more sophisticated components should only be introduced when simpler methods prove insufficient.

By Tahreem Rasul
Towards Data Science
Aug 18

Building Enterprise Agent Systems that People can Trust, Verify and Improve

The article outlines five principles that guide the successful deployment of enterprise agent systems, illustrated with a real-world example from a $100M+ company. It explains how these principles help ensure that such systems can be trusted, verified, and improved over time. The post serves as a practical guide for building reliable agent-based solutions in production environments.

By Sheila Teo
Towards Data Science
Jul 24

Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship

Enterprise Document Intelligence [Vol. 1 #8quater] - Two angles on the cascade, cost and a validation loop, backed by a real sweep of twenty local models against a hosted flagship The post Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship appeared first on Towards Data Science .

By Kezhan Shi
arXiv AI
Sep 12

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

The paper introduces a synthetic data generator that creates fully consistent, fictional enterprises—complete with workforce, customers, sales, support, and communication records—without relying on any real dataset. It validates realism through a five‑axis scorecard, an adversarial detector, and soundness checks, achieving a mean realism score of 99.1 across 23 generated companies. A second generator produces relational databases from business questions, ensuring qualifying rows and exact labels, and is available as a hosted service and containerized simulators.

By Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary, Omer Niv
arXiv Computer Vision
Aug 25

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

The paper introduces CORA (Counterfactual, Observable Redundancy Audit), a protocol for auditing website redundancy by measuring repetition load, normal-use tax, and failure-domain recovery reserve. Each audit run records screenshots, stable element identities, and task traces, while a versioned vision‑language model generates annotations that are validated and released only if they meet calibrated criteria. Experiments on a transparent mechanistic testbed show that CORA’s factorized representation separates reserve from normal-use tax and predicts perturbed success more accurately than scalar-load baselines, but it withholds automated scores when instruments fail to meet release requirements, indicating that CORA is an auditable candidate procedure for the studied benchmark rather than a universal standard.

By Ge Kong, Yongtong Cao
Towards Data Science
Aug 27

I Trained Six Models for Fraud Detection, and the Best One Isn't in Production

The article recounts a final‑year project in which the author trained six different models for fraud detection. It highlights the discrepancy between the model that performed best on evaluation metrics and the one that was ultimately chosen for production. The piece reflects on how real‑world constraints can override purely statistical performance.

By Benjamin Nweke