Towards Data Science

When the Correct Answer Is Nothing, What Does Your Pipeline Return?

The article discusses how the reliability mechanisms added to large language model (LLM) pipelines can lead to confident but incorrect outputs, especially when the correct answer is absent. It examines the behavior of pipelines in such scenarios and highlights the paradox where safeguards intended to improve accuracy may actually reinforce errors. The piece underscores the importance of understanding pipeline responses when faced with missing or ambiguous information.

Towards Data Science
Aug 31

Your LLM Can Return Perfect JSON and Still Be Wrong

The article discusses insights gained from a deeper examination of Structured Outputs when dealing with messy, incomplete data. It highlights that even when a large language model returns perfectly formatted JSON, the content can still be incorrect. The author reflects on the implications of this observation for data science practices.

By Benjamin Nweke
Towards Data Science
Aug 20

How to Fine-Tune an LLM: An End-to-End Guide

The article "How to Fine-Tune an LLM: An End-to-End Guide" offers a practical, hands‑on walkthrough for fine‑tuning large language models in real‑world scenarios. It covers the entire process from data preparation to deployment, providing readers with actionable steps to adapt LLMs to specific tasks. The guide is aimed at practitioners looking to implement fine‑tuning in a structured, end‑to‑end manner.

By Sam Black
Towards Data Science
Sep 1

Your JSON Is Valid but Your Data Is Wrong: Five Failure Modes LLM Structured Outputs Won't Catch

The article discusses five failure modes that can slip through constrained decoding in large language models, explaining why these errors are not detected by schema validators. It highlights that even when JSON output is syntactically valid, the underlying data can still be incorrect. The post serves as a warning that relying solely on schema validation is insufficient for ensuring correct structured outputs from LLMs.

By Mostafa Ibrahim
Towards Data Science
Aug 19

How to Scale an Integration Pipeline Without Breaking Correctness

The article describes a real‑world case of scaling an enterprise integration pipeline from 500 to 8,000 events per second. It emphasizes that during this throughput increase, two correctness guarantees were strictly maintained and never compromised. The post illustrates how to achieve high performance while preserving essential data integrity constraints.

By Yuelin Ou