arXiv Machine Learning By Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin, Gagan Agrawal, Wei Niu

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Read the original on arXiv Machine Learning →

EAServe introduces an encode-aware disaggregated serving framework for multimodal large language models (MLLMs), restructuring the traditional Prefill-Decode pipeline into a three-stage Encode-Prefill-Decode (EPD) system. By treating Encode as the control point, EAServe coordinates load‑adaptive micro‑batching, rate‑controlled offloading to prefill workers, and dynamic SM partitioning to balance GPU utilization across stages. Its Hybrid Auto Selection (HAS) layer optimizes GPU allocation, encode batch size, and offload ratio using capacity profiling and Bayesian optimization, achieving up to 4.3× higher goodput compared to NVIDIA Dynamo and 1.7× higher than vLLM on various MLLM architectures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral
arXiv Machine Learning
Sep 14

Efficient Vision-Language-Action Management and Serving for Robot Factories

Robion is a new serving and management system designed to run Vision‑Language‑Action (VLA) models on multi‑GPU edge servers for robot factories. It splits the VLM and ADiT stages within a single GPU, shares streams across multiple models, and prioritizes requests by remaining SLO time, enabling high robot load while meeting strict latency requirements. In experiments, Robion achieves 6.7× higher robot load than vLLM‑Omni and 1.5× higher than a monolithic pipeline, and can serve 64 robots on a 4‑GPU server with 98% SLO attainment.

By Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula
Hugging Face Trending Papers
Aug 4

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization.