Remote VAEs for decoding with Inference Endpoints 🤗
Related stories
Transformer-based Encoder-Decoder Models
Ettin Suite: SoTA Paired Encoders and Decoders
Low-Latency Task-Oriented Image Transmission with Opportunistic Spectrum Access
arXiv:2607. 01921v1 Announce Type: cross Abstract: Communication systems designed for reliable data reconstruction, rather than task-oriented communication, typically rely on separate source and channel coding and incur high latency under limited spectrum availability and fading channels.
On the quantitative analysis of decoder-based generative models
$\mathbf{\lambda}$-VAE: Variance Equalization for Posterior Collapse
arXiv:2607. 05531v1 Announce Type: new Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a failure mode in which the approximate posterior converges to the prior, rendering the latent code uninformative.
Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds
arXiv:2607. 28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments.
WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing
WhiteMatter introduces a novel architecture for Transformers that connects every attention layer to representations from all layers of each past token, allowing connection weights to vary across consumer layers and adapt to the source token. The design uses a router to mix the $L$ layer states of each token into $k$ KV channels, which are cached for subsequent tokens; each consumer layer attends to one channel. Experiments show that WhiteMatter outperforms a vanilla Transformer with 50% more layers and maintains most of this advantage even when the KV-cache is compressed by 50%.
Speculative Decoding for 2x Faster Whisper Inference
JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications
arXiv:2606. 12858v1 Announce Type: cross Abstract: Conventional communication systems, including both separation-based coding and learning-based joint source-channel coding (JSCC), are typically designed under Shannon's rate-distortion theory.
Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis
arXiv:2506. 00849v2 Announce Type: replace Abstract: Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure.
Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality
The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.