Towards Data Science

My Fall-Detection Model Scored 94%, and It Was Lying to Me

How a single evaluation choice inflated my results by 25 points, and what rebuilding honestly taught me about ML systems people might depend on The post My Fall-Detection Model Scored 94%, and It Was Lying to Me appeared first on Towards Data Science .

Towards Data Science
Aug 14

My Model Was Cheating on Its Own Test

A preprocessing pipeline let my car price model peek at the test set before the exam, and the twelve points of R squared it cheated its way to The post My Model Was Cheating on Its Own Test appeared first on Towards Data Science .

By Abdullahi Dattijo
Towards Data Science
Aug 27

I Trained Six Models for Fraud Detection, and the Best One Isn't in Production

The article recounts a final‑year project in which the author trained six different models for fraud detection. It highlights the discrepancy between the model that performed best on evaluation metrics and the one that was ultimately chosen for production. The piece reflects on how real‑world constraints can override purely statistical performance.

By Benjamin Nweke
Towards Data Science
Sep 3

My Model Worked Perfectly. Then I Tried to Make It Useful.

The article describes how to deploy a trained churn classifier as a FastAPI service so that other software can call it. It focuses on the practical steps needed to transform a model that performs well in isolation into a usable, callable API. The post is aimed at readers who want to make their machine‑learning models accessible in real-world applications.

By Ibrahim Salami
Towards Data Science
Aug 20

The LLM Judge That Kept Agreeing With Itself

The article recounts a production incident where a large language model (LLM) was used to evaluate the outputs of another LLM, and the judging model consistently agreed with itself. It explores the implications of relying on one model to assess another’s work, highlighting the potential pitfalls of such an approach. The narrative offers lessons on the limits of trusting automated evaluation systems in real‑world deployments.

By Priyansh Bhardwaj
Towards Data Science
Aug 26

How Does a RAG Reranker Really Work?

The article "How Does a RAG Reranker Really Work?" explores the inner workings of Retrieval-Augmented Generation (RAG) rerankers, focusing on how data scientists explain the model’s operations behind the scenes. It discusses the impact of these insights on architecture decisions within enterprise document intelligence, specifically in the context of Enterprise Document Intelligence Vol.1 #2D. The piece highlights the importance of transparent model explanations for effective enterprise RAG implementation.

By Kezhan Shi
arXiv Machine Learning
Jun 16

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

arXiv:2606. 15127v1 Announce Type: new Abstract: Reasoning models are increasingly used in settings where the final answer is not the only object of review: educational tools may show students intermediate steps, decision-support systems may require human oversight, and audit workflows may inspect traces for misleading or biased input.

By Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong, Yanhang Li, Zhichao Fan, Zexin Zhuang, Wenlong Dong, Zhiyuan Zheng, Hrishikesh Paranjape, Abhishek Mandal, Johnny R. Zhang