arXiv Computer Vision

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.

arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv AI
Sep 2

Superposed Latent Autoencoder

The paper introduces the Superposed Latent Autoencoder (SLAE), a method that stores multiple wide latent representations together by superposing them into a single memory tensor using learned codes and randomized keys. SLAE eliminates the need for tight dimensional bottlenecks, achieving up to 56% lower reconstruction error on datasets such as CIFAR-10/100 and SVHN while maintaining the same storage budget. The approach also boosts downstream classification performance by up to 16.79 percentage points, demonstrating that wide representations can be effectively compressed through structured interference rather than dimensional reduction.

By Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina, Tajana Rosing
arXiv Computer Vision
Aug 25

HP-UniIF: Hierarchical Prompt Learning for Unified Image Fusion

HP-UniIF is a unified vision framework that uses diffusion priors and a depth‑wise hierarchical conditional modulation strategy to support heterogeneous image fusion, visual restoration, and downstream perception tasks. The framework introduces task prompt modulation at bottleneck layers, a degradation prompt router at shallow layers, and an application prompt bank at decoding stages to decouple and adapt to different objectives. Experiments across multiple fusion tasks, degradations, and downstream applications show that HP‑UniIF achieves superior performance while maintaining visually faithful results and task‑relevant semantics.

By Xingxin Xu, Siqi Zhao, Xin Li, Xinjie Yao, Yiming Sun, Pengfei Zhu