Hugging Face Blog Dec 3, 2024 Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
Sebastian Raschka Jun 17, 2025 Understanding and Coding the KV Cache in LLMs from Scratch KV caches are one of the most critical techniques for efficient inference in LLMs in production. By Sebastian Raschka, PhD
Hugging Face Blog Apr 16, 2025 Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
Sebastian Raschka May 16 Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs By Sebastian Raschka, PhD