Towards Data Science

Kimi K3’s 1M Token Context Window vs. RAG: Cost, Latency and Answer Quality

The article presents a controlled comparison between a top‑5 Retrieval‑Augmented Generation (RAG) pipeline and a single 127,000‑token prompt using the same 12 questions, system prompt, and model. Both approaches were graded blind on correctness, completeness, and grounding. The study evaluates cost, latency, and answer quality for each method.

arXiv AI
Aug 19

Token Optimization and Context Window Management in Multi-Agent AI Workflows

The paper "Token Optimization and Context Window Management in Multi‑Agent AI Workflows" introduces a practitioner framework that reduces token usage and latency in multi‑agent AI systems. It outlines six patterns—context stratification, fetch‑once/process‑locally architecture, schema‑contracted prompts, token‑aware fallback chains, semantic caching, and inter‑agent communication compression—and reports a 60‑70% token reduction and a 61‑116 second cold‑load latency improvement in production. A controlled study on relevance‑contrast context shows that mixing high‑ and low‑relevance items in prompts can improve relevance accuracy by up to +0.084. whyItMatters":"The work provides concrete, repeatable engineering patterns that bridge research and production, enabling faster, cheaper, and more reliable AI workflows."

By Dvir Shamay
Towards Data Science
Jun 18

Dispatching the Parsed RAG Question: Chunk Strategy, Model Tier, Activations, Audit

Enterprise Document Intelligence [Vol. 1 #6c] - The decisions the parser makes on top of the user string, using the document’s profile: dispatch, activations, full schema, three approaches to deciding what fires, the audit _meta block, and a broker-corpus walkthrough The post Dispatching the Parsed RAG Question: Chunk Strategy, Model Tier, Activations, Audit appeared first on Towards Data Science .

By angela shi