← Back to all news
arXiv Machine Learning August 24, 2026 By Yujie Zhang, Shivam Aggarwal, Tulika Mitra

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 29

SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models

arXiv:2607. 24787v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited accelerator memory.

By Jinwei Kong, Runqi Meng, Fanyi Wang, Wentao Qiu, Haotian Hu, Yongjian Zhou, Zhenhua Ge
efficiencybenchmarks
More like this →
arXiv AI
Aug 13

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

arXiv:2608. 11688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency.

By Alish Kanani, Layan Badawi, Umit Y. Ogras
efficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 2

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving

arXiv:2606. 00735v1 Announce Type: cross Abstract: In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execution, where the slowest GPU determines layer latency.

By Seokjin Go, Marko Scrbak, Ephrem Wu, Srilatha Manne, Divya Mahajan
llmsefficiency
More like this →
arXiv AI
Jul 14

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

arXiv:2607. 10183v1 Announce Type: cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory.

By Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li
llmsefficiency
More like this →
arXiv AI
Aug 11

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

arXiv:2608. 07964v1 Announce Type: cross Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models.

By Yize Wu, Ke Gao, Ling Li, Yanjun Wu
More like this →
arXiv AI
Jul 3

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.

By Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty
fine-tuningefficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea