Hugging Face Blog Apr 29, 2025 Introducing AutoRound: Intelβs Advanced Quantization for LLMs and VLMs
Hugging Face Blog May 24, 2023 Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
Sebastian Raschka Jun 17, 2025 Understanding and Coding the KV Cache in LLMs from Scratch KV caches are one of the most critical techniques for efficient inference in LLMs in production. By Sebastian Raschka, PhD
Sebastian Raschka May 16 Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs By Sebastian Raschka, PhD