← Back to all news
arXiv Machine Learning September 22, 2026 By Yifan Sui, Hao Wang, Hanfei Yu, Kaiqiang Xu, Yitao Hu, Chen Chen, Jianxun Li, Kai Chen

ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 31

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

arXiv:2607. 26566v1 Announce Type: cross Abstract: Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently.

By Xiaoxiao Jiang, Suyi Li, Sheng Yao, Tianyu Feng, Lingyun Yang, Dapeng Nie, Haoran Yang, Wei Wang
diffusionsafety
More like this →
arXiv Machine Learning
Aug 3

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

arXiv:2607. 28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak.

By Jiaxuan Chen, Jianshu She, Ye Yuan, Rajat Ghosh, Karan Gupta, Qirong Ho, Xue Liu, Oana Balmau
llmsfine-tuningefficiency
More like this →
Hugging Face Trending Papers
Jul 29

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together.

diffusionsafety
More like this →
arXiv AI
Jul 21

Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs

arXiv:2607. 17181v1 Announce Type: cross Abstract: Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently.

By Utopia Meng, Unicornt Zhao, Derek Li, Goalen Gao, Frank Du
llmsagents
More like this →
arXiv AI
Jun 2

Lodestar: An Online-Learning LLM Inference Router

arXiv:2606. 00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization.

By Gangmuk Lim, Wanyu Zhao, Brighten Godfrey, Jiaxin Shan, Le Xu, Liguang Xie
llmsbenchmarks
More like this →
arXiv Machine Learning
Aug 18

Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Serving

arXiv:2608. 16336v1 Announce Type: cross Abstract: Modern LLM serving deployments must simultaneously satisfy heterogeneous service-level objectives (SLOs) across a diverse population of user tiers, ranging from latency-critical API calls to background batch processing.

By Anders Vestrum, Arya Raeesi, Hanna Roed
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea