Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs
Understanding How DeepSeek's Flagship Open-Weight Models Evolved
A Round Up And Comparison of 10 Open-Weight LLM Releases in Spring 2026
KV caches are one of the most critical techniques for efficient inference in LLMs in production.
Beat the 8GB VRAM limit. Learn how to run three different LLMs on a single 8GB GPU using C++ layer multiplexing and admission control.
arXiv:2606. 09879v1 Announce Type: new Abstract: This study addresses on-device inference bottlenecks of Transformer models on Tenstorrent's Tensix architecture and proposes an operator fusion strategy that enhances data locality.
arXiv:2605. 25451v2 Announce Type: replace Abstract: Training multimodal large language models (MLLMs) is challenged by both model and data heterogeneity.
arXiv:2607. 26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs.
arXiv:2608. 05303v1 Announce Type: cross Abstract: On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications.
arXiv:2607. 06839v1 Announce Type: new Abstract: Existing NAS benchmarks (e.