arXiv:2606. 05981v1 Announce Type: cross Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1-step distilled student, the text encoder becomes the critical path.
By Yoshiyuki Ootani
arXiv:2604. 16514v5 Announce Type: replace-cross Abstract: Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck.
By Baoyou Chen, Hanchen Xia, Peng Tu, Haojun Shi, Liwei Zhang, Yuxuan Yao, Weihao Yuan, Siyu Zhu
arXiv:2609.37831v1 Announce Type: new
Abstract: Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise gene...
By Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, Zhibo Chen
EAServe introduces an encode-aware disaggregated serving framework for multimodal large language models (MLLMs), restructuring the traditional Prefill-Decode pipeline into a three-stage Encode-Prefill-Decode (EPD) system. By treating Encode as the control point, EAServe coordinates load‑adaptive micro‑batching, rate‑controlled offloading to prefill workers, and dynamic SM partitioning to balance GPU utilization across stages. Its Hybrid Auto Selection (HAS) layer optimizes GPU allocation, encode batch size, and offload ratio using capacity profiling and Bayesian optimization, achieving up to 4.3× higher goodput compared to NVIDIA Dynamo and 1.7× higher than vLLM on various MLLM architectures.
By Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin, Gagan Agrawal, Wei Niu
arXiv:2512.23709v3 Announce Type: replace
Abstract: Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due...
By Hau-Shiang Shiu, Chin-Yang Lin, Zhixiang Wang, Chi-Wei Hsiao, Po-Fan Yu, Yu-Chih Chen, Yu-Lun Liu
arXiv:2607. 14898v1 Announce Type: cross Abstract: Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video.
By Minguk Kang, Suha Kwak