Take What You Need: Flexible Multi-Task Semantic Communications with Channel Adaptation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to jointly perform heterogeneous classification and regression tasks on the Cityscapes dataset, using an InfoMax principle to handle mixed discrete and continuous semantic variables. It compares the new framework against independent single‑task training, conventional task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom, and studies how the capacity of the common unit affects joint task performance. Extensive evaluations show that CMT‑SemCom outperforms all benchmarks.
The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to handle heterogeneous classification and regression tasks on the Cityscapes dataset. By incorporating the InfoMax principle, the system accommodates mixed discrete and continuous semantic variables, and it is benchmarked against single‑task training, task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom. Experiments show that CMT‑SemCom outperforms all baselines and provide insights into how the common unit capacity affects joint task performance.
The paper proposes a task-oriented semantic feature transmission framework for satellite remote sensing over low‑signal‑to‑noise ratio (SNR) channels. Instead of reconstructing images first, it directly transmits semantic features extracted by a multitask‑pretrained backbone, using a lightweight channel adaptation module to reduce bandwidth and a feature restorer to recover task‑relevant structure after channel corruption. Experiments on scene classification and object detection under additive white Gaussian noise show consistent improvements over reconstruction‑oriented joint source‑channel coding baselines, especially in the low‑SNR regime.
ShallowStream is a framework for streaming video understanding that uses the shallow layers of a multimodal large language model (MLLM) to encode frames and build a lightweight index. During streaming, it maintains an always‑on index via the KV cache of shallow layers, and at query time it scores context frames using shallow‑layer attention and selects diverse evidence for answering. The approach matches the performance of leading streaming methods while cutting per‑frame prefill latency and 10‑second end‑to‑end latency by up to 52.1× and 11.9×, respectively.
The paper introduces a token‑oriented semantic communication framework that transmits only task‑relevant image latents instead of full token embeddings, reducing communication cost and improving interoperability. It leverages a spatial alignment between vision transformer patch tokens and learned image compression latents, enabling token‑level relevance estimation and selective transmission. Experiments on ImageNet demonstrate a superior rate–accuracy trade‑off compared to existing semantic communication methods and hand‑crafted codecs.
arXiv:2606. 10277v1 Announce Type: new Abstract: Though wireless foundation models (WFMs) have shown strong potential in learning universal channel representations, their adaptation to various downstream tasks remains constrained by existing paradigms.