arXiv Machine Learning

Channel-Adaptive Region Adjacency Graph Carriers for Semantic Image Communication

arXiv Machine Learning
Aug 27

Token-Oriented Semantic Communication with Pretrained Vision Transformers

The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.

By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim
arXiv Computer Vision
3d ago

Rethinking Generative Image Compression at Extremely Low Bitrates

The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.

By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv AI
Jun 2

Channel-wise Vector Quantization

arXiv:2605. 26089v2 Announce Type: replace-cross Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens.

By Wei Song, Tianhang Wang, Yitong Chen, Tong Zhang, Zuxuan Wu, Min Li, Jiaqi Wang, Kaicheng Yu
Hugging Face Trending Papers
Jul 27

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency.

arXiv AI
Sep 10

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

The paper introduces FMSGOC, a framework that leverages visual‑linguistic foundation models to improve semantic and goal‑oriented communication for 6G. By transmitting a sparse set of semantic anchors and using a pretrained diffusion model for masked completion, it reduces overfitting and achieves high rate efficiency, reaching 0.039 BPP while maintaining strong semantic fidelity and robustness on unseen data.

By Boliang Liu, Wint Yi Poe, Riccardo Trivisonno, Giuseppe Caire