Building a Data Lakehouse with DuckDB and DuckLake
Related stories
Running SQL Concurrently Across Three Remote DuckDB Servers with Quack
A small experiment in remote SQL execution The post Running SQL Concurrently Across Three Remote DuckDB Servers with Quack appeared first on Towards Data Science .
I Deployed My Data Pipeline to AWS. Then Everything That Was “Local” Broke.
The article recounts the author's experience of moving a Dockerized data pipeline from a local laptop to AWS, highlighting the challenges that arose when the environment changed. It explores lessons learned about container behavior, networking intricacies, and hidden assumptions that were previously taken for granted in a local setup. The post serves as a practical guide for developers facing similar transitions to cloud infrastructure.
Avoiding Entity Key Drift in a Data Lake: Step 1, Normalization
This is the opening piece of a four-part deep dive series, on building a high-frequency streaming pipeline against a live public API. The data source is openSenseMap, a citizen-science IoT network used for climate research, mostly in Germany.
Stop Choosing Between Local and Cloud LLMs: A Field Guide to Hybrid Patterns
A hands-on walkthrough of a hybrid local-cloud workflow using Gemma 4 and GPT-5. 4, with reasoning and structured outputs The post Stop Choosing Between Local and Cloud LLMs: A Field Guide to Hybrid Patterns appeared first on Towards Data Science .
Connecting My LangGraph AI Agent to Postgres
The article explains how to connect a LangGraph AI agent to a Postgres database. It covers running the backend locally using Docker and deploying it in the cloud. The guide provides practical steps for setting up the database connection and managing the agent’s data storage.
The Medallion Data Architecture: An Introduction
A practical guide to Bronze, Silver and Gold, with a working Python and DuckDB example The post The Medallion Data Architecture: An Introduction appeared first on Towards Data Science .
Parse PDFs for RAG Locally with Docling: Rich Tables, No Cloud Upload
Enterprise Document Intelligence [Vol. 1 #5ter] - Table cells, OCR, captions, headings: cloud-grade structure, running on your own machine.
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Building a production-ready RSS pipeline with Python, Docker, PostgreSQL, and Kestra The post I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer appeared first on Towards Data Science .
What Are the Possibilities to Build Date Tables in Self-Service Environments?
For years, I created date tables with DAX code whenever I didn’t have a way to create them upstream of the data flow. Now I've realised there's another way to do it.
PySpark for Beginners: Building Intermediate-Level Skills
A practical next step into partitions, shuffles, joins, caching, and execution plans. The post PySpark for Beginners: Building Intermediate-Level Skills appeared first on Towards Data Science .
alchemy-utils 0.1a1
Release: alchemy-utils 0. 1a1 Performance boost for DuckDB exports and CSV imports, see here .