← Back to all news
arXiv Computer Vision October 1, 2026 By Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen

DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 7

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

arXiv:2511. 04805v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models have shown strong potential in scaling language models efficiently by activating only a small subset of experts per input.

By Yushu Zhao, Zheng Wang, Minjia Zhang
llmsbenchmarks
More like this →
arXiv AI
Jul 1

Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

arXiv:2606. 31903v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation.

By Zhaoyang Luo, Runmin Dong, Miao Yang, Fan Wei, Yushan Lai, Bin Luo, Haohuan Fu
llmsmultimodalbenchmarks
More like this →
arXiv Machine Learning
Jun 11

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models

arXiv:2605. 25820v2 Announce Type: replace Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel.

By Yulin Yuan, Hongshuo Zhao, Xiangming Meng
llmsdiffusionmultimodalbenchmarks
More like this →
arXiv Computer Vision
Sep 2

Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference

arXiv:2609.01200v1 Announce Type: new Abstract: When the visual encoder and the language decoder of a vision-language model (VLM) run on different compute nodes, the intermediate visual-token embeddi...

By Reza Heidari, Hamed R. Tavakoli, Juho Kannala
llmsragnlpefficiencymultimodal
More like this →
arXiv AI
Sep 15

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

arXiv:2609.15131v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most method...

By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang
llmsreinforcement-learningmultimodal
More like this →
arXiv Machine Learning
Aug 26

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

arXiv:2608.24763v1 Announce Type: cross Abstract: Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedu...

By Muhammad Asad Ali, Umar Khan, Nadia Robertini, Didier Stricker
llmsmultimodalbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea