The paper introduces MIRC, an overfitted image codec that quantizes and entropy‑codes all components—including latents, synthesis network, and entropy models—within a single end‑to‑end rate‑distortion framework inspired by NVRC. It adds a multi‑scale representation with cross‑stage parameter sharing to capture cross‑scale redundancy, yielding a 10.5 % BD‑rate saving over VVC on the CLIC2020 professional set. MIRC offers multiple configurations ranging from 1.2 to 2.9 kMAC per pixel, allowing decoding complexity to be tuned to deployment needs.
By Tianhao Peng, Ho Man Kwan, Fan Zhang, Shan Liu, David Bull
The paper introduces MIRC, an overfitted image codec that quantizes all components—including latents, synthesis network, and entropy models—within a single rate‑distortion objective, following the neural video representation codec NVRC. It adds a multi‑scale representation with cross‑stage parameter sharing to capture cross‑scale redundancy, achieving a 10.5% BD‑rate saving over VVC on the CLIC2020 professional validation set. MIRC also offers configurable decoding complexity ranging from 1.2 to 2.9 kMAC per pixel, allowing deployment to match specific resource budgets.
Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss, a common challenge in satellite and emergency communications.
arXiv:2608.22996v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution s...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv:2601. 22002v5 Announce Type: replace Abstract: Transformers achieve superior performance on many tasks, but impose heavy compute and memory requirements during inference.
By Anderson de Andrade, Alon Harell, Ivan V. Baji\'c
The paper introduces an extension of the Decoupled Vision Transformer (DC‑ViT) to handle multi‑channel imaging (MCI) data, where each channel carries a distinct semantic signal. By tokenizing each channel separately and then decoupling intra‑channel from inter‑channel updates, the model preserves channel‑specific features. The authors further enable independent per‑channel masking by solving a linear assignment between retained patches, allowing masked training without restricting visible tokens. Experiments on fluorescence microscopy, imaging mass cytometry, and satellite imaging datasets demonstrate that this approach outperforms the strongest Multi‑Channel Vision Transformer baselines on both classification and dense‑prediction tasks.
By Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito