Towards Data Science

Are Your ML Experiments a Mess? Here’s the Fix

A hands-on guide to tracking experiments, logging models, and reproducing results with ML Flow. The post Are Your ML Experiments a Mess?

Towards Data Science
Aug 31

AgentOps Is Not MLOps: What Breaks in Your Monitoring Stack When Agents Go to Production

The article explains how the five core assumptions of MLOps monitoring are violated when agents are deployed to production, leading to inherited signals that incorrectly mark failed runs as healthy. It highlights the specific ways in which agent-based systems disrupt traditional monitoring stacks and the implications for reliability and performance. The piece serves as a warning for practitioners transitioning from MLOps to AgentOps, outlining the critical monitoring gaps that arise.

By Mostafa Ibrahim
Towards Data Science
Aug 20

How to Fine-Tune an LLM: An End-to-End Guide

The article "How to Fine-Tune an LLM: An End-to-End Guide" offers a practical, hands‑on walkthrough for fine‑tuning large language models in real‑world scenarios. It covers the entire process from data preparation to deployment, providing readers with actionable steps to adapt LLMs to specific tasks. The guide is aimed at practitioners looking to implement fine‑tuning in a structured, end‑to‑end manner.

By Sam Black
Simon Willison
4d ago

Quoting Anthropic Frontier Red Team

The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.

Towards Data Science
Sep 24

When the Correct Answer Is Nothing, What Does Your Pipeline Return?

The article discusses how the reliability mechanisms added to large language model (LLM) pipelines can lead to confident but incorrect outputs, especially when the correct answer is absent. It examines the behavior of pipelines in such scenarios and highlights the paradox where safeguards intended to improve accuracy may actually reinforce errors. The piece underscores the importance of understanding pipeline responses when faced with missing or ambiguous information.

By Hubert García Gordon
Towards Data Science
Aug 31

Your LLM Can Return Perfect JSON and Still Be Wrong

The article discusses insights gained from a deeper examination of Structured Outputs when dealing with messy, incomplete data. It highlights that even when a large language model returns perfectly formatted JSON, the content can still be incorrect. The author reflects on the implications of this observation for data science practices.

By Benjamin Nweke
Towards Data Science
Sep 3

My Model Worked Perfectly. Then I Tried to Make It Useful.

The article describes how to deploy a trained churn classifier as a FastAPI service so that other software can call it. It focuses on the practical steps needed to transform a model that performs well in isolation into a usable, callable API. The post is aimed at readers who want to make their machine‑learning models accessible in real-world applications.

By Ibrahim Salami