Towards Data Science

My Model Was Cheating on Its Own Test

A preprocessing pipeline let my car price model peek at the test set before the exam, and the twelve points of R squared it cheated its way to The post My Model Was Cheating on Its Own Test appeared first on Towards Data Science .

Towards Data Science
Aug 27

I Trained Six Models for Fraud Detection, and the Best One Isn't in Production

The article recounts a final‑year project in which the author trained six different models for fraud detection. It highlights the discrepancy between the model that performed best on evaluation metrics and the one that was ultimately chosen for production. The piece reflects on how real‑world constraints can override purely statistical performance.

By Benjamin Nweke
Towards Data Science
Aug 20

The LLM Judge That Kept Agreeing With Itself

The article recounts a production incident where a large language model (LLM) was used to evaluate the outputs of another LLM, and the judging model consistently agreed with itself. It explores the implications of relying on one model to assess another’s work, highlighting the potential pitfalls of such an approach. The narrative offers lessons on the limits of trusting automated evaluation systems in real‑world deployments.

By Priyansh Bhardwaj
Towards Data Science
Sep 3

My Model Worked Perfectly. Then I Tried to Make It Useful.

The article describes how to deploy a trained churn classifier as a FastAPI service so that other software can call it. It focuses on the practical steps needed to transform a model that performs well in isolation into a usable, callable API. The post is aimed at readers who want to make their machine‑learning models accessible in real-world applications.

By Ibrahim Salami
Towards Data Science
Sep 22

Break Your Own RAG Pipeline Before Users Do

The article discusses a small adversarial test set designed to detect retrieval failures in Retrieval-Augmented Generation (RAG) pipelines that typical evaluation sets might miss. It emphasizes the importance of proactively testing your own RAG system to uncover hidden weaknesses before users encounter them. By using this targeted test set, developers can improve the reliability and robustness of their RAG models.

By Sara Nobrega
Towards Data Science
2d ago

Your AI Bill Is a Toll Booth. Stop Paying Twice.

The article titled "Your AI Bill Is a Toll Booth. Stop Paying Twice." discusses how users are unexpectedly paying more for AI services than anticipated, likening the experience to a toll booth where one pays twice. It highlights the unseen costs that can arise when using AI tools and urges readers to be vigilant about their expenses. The piece was first published on Towards Data Science.

By Gursimar Singh
Towards Data Science
Aug 31

Your LLM Can Return Perfect JSON and Still Be Wrong

The article discusses insights gained from a deeper examination of Structured Outputs when dealing with messy, incomplete data. It highlights that even when a large language model returns perfectly formatted JSON, the content can still be incorrect. The author reflects on the implications of this observation for data science practices.

By Benjamin Nweke