Towards Data Science

Cut an Enterprise RAG Pipeline’s Latency and Cost by Calling the LLM Less, Not by Buying a Faster Model

Enterprise Document Intelligence [Vol. 1 #9ter] - The pipeline from Article 9 calls a model at several steps to be sure it is right.

Towards Data Science
Jul 24

Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship

Enterprise Document Intelligence [Vol. 1 #8quater] - Two angles on the cascade, cost and a validation loop, backed by a real sweep of twenty local models against a hosted flagship The post Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Aug 19

Kimi K3’s 1M Token Context Window vs. RAG: Cost, Latency and Answer Quality

The article presents a controlled comparison between a top‑5 Retrieval‑Augmented Generation (RAG) pipeline and a single 127,000‑token prompt using the same 12 questions, system prompt, and model. Both approaches were graded blind on correctness, completeness, and grounding. The study evaluates cost, latency, and answer quality for each method.

By Sarah Schürch
Towards Data Science
Aug 20

How to Fine-Tune an LLM: An End-to-End Guide

The article "How to Fine-Tune an LLM: An End-to-End Guide" offers a practical, hands‑on walkthrough for fine‑tuning large language models in real‑world scenarios. It covers the entire process from data preparation to deployment, providing readers with actionable steps to adapt LLMs to specific tasks. The guide is aimed at practitioners looking to implement fine‑tuning in a structured, end‑to‑end manner.

By Sam Black
Simon Willison
Sep 20

Quoting voxium

Simon Willison describes his experience at a large company where all documentation, code, tests, PRDs, tickets, and reports are generated by Claude Code. His team is forced to ship rapidly, working long hours, yet management insists that code push is not a bottleneck, leading to frustration and a lack of meaningful reading or review. The situation highlights a reliance on AI-generated content that may undermine quality and collaboration.

Towards Data Science
May 29

RAG Is Burning Money — I Built a Cost Control Layer to Fix It

Most RAG systems are optimized for answer quality, not cost—and that blind spot gets expensive fast. In this article, I break down a production-ready cost control layer combining semantic caching, query routing, token budgeting, and circuit breaking, achieving an 85% reduction in LLM costs without sacrificing answer quality.

By Emmimal P Alexander
Towards Data Science
Aug 19

How to Scale an Integration Pipeline Without Breaking Correctness

The article describes a real‑world case of scaling an enterprise integration pipeline from 500 to 8,000 events per second. It emphasizes that during this throughput increase, two correctness guarantees were strictly maintained and never compromised. The post illustrates how to achieve high performance while preserving essential data integrity constraints.

By Yuelin Ou
Towards Data Science
Jul 22

Loop Engineering for RAG Generation: Iterate top-k One at a Time

Enterprise Document Intelligence [Vol. 1 #8bis] - Two regimes for sending retrieved candidates to the generation brick, the sufficiency signal that picks between them, and the per-question type dispatch that makes it cheap The post Loop Engineering for RAG Generation: Iterate top-k One at a Time appeared first on Towards Data Science .

By Kezhan Shi