Towards Data Science

I Deployed My Data Pipeline to AWS. Then Everything That Was “Local” Broke.

The article recounts the author's experience of moving a Dockerized data pipeline from a local laptop to AWS, highlighting the challenges that arose when the environment changed. It explores lessons learned about container behavior, networking intricacies, and hidden assumptions that were previously taken for granted in a local setup. The post serves as a practical guide for developers facing similar transitions to cloud infrastructure.

Towards Data Science
Sep 24

When the Correct Answer Is Nothing, What Does Your Pipeline Return?

The article discusses how the reliability mechanisms added to large language model (LLM) pipelines can lead to confident but incorrect outputs, especially when the correct answer is absent. It examines the behavior of pipelines in such scenarios and highlights the paradox where safeguards intended to improve accuracy may actually reinforce errors. The piece underscores the importance of understanding pipeline responses when faced with missing or ambiguous information.

By Hubert García Gordon
Towards Data Science
Aug 19

How to Scale an Integration Pipeline Without Breaking Correctness

The article describes a real‑world case of scaling an enterprise integration pipeline from 500 to 8,000 events per second. It emphasizes that during this throughput increase, two correctness guarantees were strictly maintained and never compromised. The post illustrates how to achieve high performance while preserving essential data integrity constraints.

By Yuelin Ou