Towards Data Science By angela shi

Cut an Enterprise RAG Pipeline’s Latency and Cost by Calling the LLM Less, Not by Buying a Faster Model

Read the original on Towards Data Science →

Enterprise Document Intelligence [Vol. 1 #9ter] - The pipeline from Article 9 calls a model at several steps to be sure it is right.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

Towards Data Science
Jul 24

Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship

Enterprise Document Intelligence [Vol. 1 #8quater] - Two angles on the cascade, cost and a validation loop, backed by a real sweep of twenty local models against a hosted flagship The post Loop Engineering for RAG Generation: An LLM Cascade from a Cheap Local Model Up to a Hosted Flagship appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Aug 19

Kimi K3’s 1M Token Context Window vs. RAG: Cost, Latency and Answer Quality

The article presents a controlled comparison between a top‑5 Retrieval‑Augmented Generation (RAG) pipeline and a single 127,000‑token prompt using the same 12 questions, system prompt, and model. Both approaches were graded blind on correctness, completeness, and grounding. The study evaluates cost, latency, and answer quality for each method.

By Sarah Schürch
Towards Data Science
Aug 20

How to Fine-Tune an LLM: An End-to-End Guide

The article "How to Fine-Tune an LLM: An End-to-End Guide" offers a practical, hands‑on walkthrough for fine‑tuning large language models in real‑world scenarios. It covers the entire process from data preparation to deployment, providing readers with actionable steps to adapt LLMs to specific tasks. The guide is aimed at practitioners looking to implement fine‑tuning in a structured, end‑to‑end manner.

By Sam Black
Simon Willison
Sep 20

Quoting voxium

Simon Willison describes his experience at a large company where all documentation, code, tests, PRDs, tickets, and reports are generated by Claude Code. His team is forced to ship rapidly, working long hours, yet management insists that code push is not a bottleneck, leading to frustration and a lack of meaningful reading or review. The situation highlights a reliance on AI-generated content that may undermine quality and collaboration.

Towards Data Science
May 29

RAG Is Burning Money — I Built a Cost Control Layer to Fix It

Most RAG systems are optimized for answer quality, not cost—and that blind spot gets expensive fast. In this article, I break down a production-ready cost control layer combining semantic caching, query routing, token budgeting, and circuit breaking, achieving an 85% reduction in LLM costs without sacrificing answer quality.

By Emmimal P Alexander