Making thousands of open LLMs bloom in the Vertex AI Model Garden
Related stories
Optimizing your LLM in production
Optimization story: Bloom inference
How To Build Your Own LLM Runtime From Scratch
If you have ever wanted to actually build an LLM inference runtime yourself — pack your own weights, own every barrier, capture your own CUDA graphs — this is what that journey looks like on an H100. A step-by-step tour of a small runtime called annotated-llm-runtime, and the three bugs that produced most of the annotations.
Understanding and Coding the KV Cache in LLMs from Scratch
KV caches are one of the most critical techniques for efficient inference in LLMs in production.
Apertus LLM Family Expansion via Distillation and Quantization
arXiv:2605. 29128v2 Announce Type: replace Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints.
My Workflow for Understanding LLM Architectures
A learning-oriented workflow for understanding new open-weight model releases
Deploy Meta Llama 3.1 405B on Google Cloud Vertex AI
A Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026
A Round Up And Comparison of 10 Open-Weight LLM Releases in Spring 2026
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs



