arXiv Computer Vision

Take What You Need: Flexible Multi-Task Semantic Communications with Channel Adaptation

arXiv Machine Learning
Sep 4

Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks

The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to jointly perform heterogeneous classification and regression tasks on the Cityscapes dataset, using an InfoMax principle to handle mixed discrete and continuous semantic variables. It compares the new framework against independent single‑task training, conventional task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom, and studies how the capacity of the common unit affects joint task performance. Extensive evaluations show that CMT‑SemCom outperforms all benchmarks.

By Ahmad Halimi Razlighi, Mohammad Siddiqur Rahman, Maximilian H. V. Tillmann, Edgar Beck, Armin Dekorsy
Hugging Face Trending Papers
Sep 3

Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks

The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to handle heterogeneous classification and regression tasks on the Cityscapes dataset. By incorporating the InfoMax principle, the system accommodates mixed discrete and continuous semantic variables, and it is benchmarked against single‑task training, task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom. Experiments show that CMT‑SemCom outperforms all baselines and provide insights into how the common unit capacity affects joint task performance.

arXiv Machine Learning
Sep 18

Task-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels

The paper proposes a task-oriented semantic feature transmission framework for satellite remote sensing over low‑signal‑to‑noise ratio (SNR) channels. Instead of reconstructing images first, it directly transmits semantic features extracted by a multitask‑pretrained backbone, using a lightweight channel adaptation module to reduce bandwidth and a feature restorer to recover task‑relevant structure after channel corruption. Experiments on scene classification and object detection under additive white Gaussian noise show consistent improvements over reconstruction‑oriented joint source‑channel coding baselines, especially in the low‑SNR regime.

By Shuoyuan Sun, Hongyu Wang, Mugen Peng, Wenjia Xu
arXiv Computation and Language
Sep 3

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

ShallowStream is a framework for streaming video understanding that uses the shallow layers of a multimodal large language model (MLLM) to encode frames and build a lightweight index. During streaming, it maintains an always‑on index via the KV cache of shallow layers, and at query time it scores context frames using shallow‑layer attention and selects diverse evidence for answering. The approach matches the performance of leading streaming methods while cutting per‑frame prefill latency and 10‑second end‑to‑end latency by up to 52.1× and 11.9×, respectively.

By Jitai Hao, Ke Yang, Qiang Huang, Jun Yu
arXiv Machine Learning
Aug 27

Token-Oriented Semantic Communication with Pretrained Vision Transformers

The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.

By Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim
arXiv Machine Learning
Jun 10

A Unified Adaptive Feature Composition Framework for Multi-Task Generalization in Wireless Foundation Models

arXiv:2606. 10277v1 Announce Type: new Abstract: Though wireless foundation models (WFMs) have shown strong potential in learning universal channel representations, their adaptation to various downstream tasks remains constrained by existing paradigms.

By Yuxuan Shi, Tingting Yang, Kangning Ma, Liwen Jing, Yuwei Wang, Mengfan Zheng, Li Sun
arXiv Machine Learning
Sep 16

Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

The paper introduces a semantic‑aware multi‑level neural video codec designed for low‑latency, task‑oriented video transmission over unreliable channels. It builds on the real‑time DCVC‑RT codec by partitioning encoded representations into packets of varying semantic and feature importance, assigning them to priority streams, and employing an error‑resilient entropy model that removes inter‑packet dependencies. Experiments demonstrate that this framework improves robustness against packet erasures, achieving graceful degradation in less important regions while preserving task‑relevant visual content.

By Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan