arXiv Computer Vision By Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

Read the original on arXiv Computer Vision →

The paper introduces Q-TOFC, a query‑guided task‑oriented visual feature compression method that uses residual vector quantization to encode merged features as compact codebook index sequences. By incorporating query relevance into feature aggregation and adding a quantization error compensation adapter, Q‑TOFC reduces visual payload by 53.6% compared to previous TOFC while preserving task performance. Experiments across seven multimodal benchmarks and latency tests confirm its effectiveness under bandwidth‑constrained uplinks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 23

Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints. This paper reviews recent advances in efficient vision-language and multimodal large language models, covering visual token compression, video token management, KV-cache optimization, Mixture-of-Experts (MoE) routing, low-bit quantization, edge deployment, and hardware-aware benchmarking.

arXiv AI
Jul 24

Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

arXiv:2607. 20981v1 Announce Type: new Abstract: Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints.

By Jay Gor, Karm Dave, Akshita Abrol, Rajesh Gupta, Sudeep Tanwar, Zhengkui Wang
arXiv Machine Learning
Sep 11

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

The paper introduces EMMI, a framework that enables communication‑efficient inference of multimodal large language models (MLLMs) on edge devices. EMMI encodes each sensor modality separately, fuses the representations, and compresses them into a compact latent vector that is transmitted to a server for high‑capacity reasoning. Experiments on a multimodal benchmark show that EMMI can cut the communication payload by 32× while keeping accuracy comparable, achieving up to a 3.4× reduction in end‑to‑end inference latency under bandwidth‑constrained conditions.

By Motahare Mounesan, Irfan Khan