TGI Multi-LoRA: Deploy Once, Serve 30 Models
Related stories
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
arXiv:2510. 13537v2 Announce Type: replace-cross Abstract: On-device deployment of Large Language Models (LLMs) frequently leverages Low-Rank Adapters (LoRAs) to support diverse downstream tasks under tight resource constraints.
PLoRA: Efficient Concurrent LoRA Training for Large Language Models
arXiv:2508. 02932v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has gained popularity as a fine-tuning approach for Large Language Models (LLMs) due to its low resource requirements and good performance.
Making thousands of open LLMs bloom in the Vertex AI Model Garden
Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs
arXiv:2607. 17181v1 Announce Type: cross Abstract: Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently.
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
Deploy models on AWS Inferentia2 from Hugging Face
Best practices for deploying language models
Cohere, OpenAI, and AI21 Labs have developed a preliminary set of best practices applicable to any organization developing or deploying large language models.
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
arXiv:2607. 02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window.