arXiv Computer Vision

Sphere Encoder 2

arXiv Computer Vision
Sep 22

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

The paper introduces the Geometry‑Native Autoencoder (GAE), a compact latent space that can be decoded into appearance, depth, camera parameters, and point maps, enabling 3D‑consistent world generation. By reparameterizing a geometry foundation model’s features, GAE replaces traditional appearance‑centric latents and improves visual quality and 3D coherence, achieving significant reductions in FVD and camera‑trajectory error on benchmark datasets. The work demonstrates that a geometry‑native latent space can serve as a shared interface between perception and generation models.

By Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
arXiv Machine Learning
Sep 24

On the Diffusibility of High-Dimensional Latents

The paper investigates how fine‑tuning pretrained visual encoders for faithful image reconstruction affects diffusion models that operate in the resulting latent space. It finds that such fine‑tuning reduces the effective dimensionality of the latent representation, causing standard velocity‑prediction flow‑matching to fit noise outside the low‑dimensional signal manifold and making optimization inefficient. Consequently, the authors propose using a clean‑data ($oldsymbol{x}_{0}$) parameterization, which focuses learning on the signal manifold and consistently improves text‑to‑image generation across multiple strong‑reconstruction encoders.

By Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li
arXiv Machine Learning
Sep 3

Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation

LF-MultiDiffusion is a training‑free method for generating spherical panoramas that extends MultiDiffusion by adding linear projections between target and reference image spaces. It reformulates latent aggregation as a regularized least‑squares problem and solves it with a Krylov‑based iterative solver during denoising, enabling denser and more natural mappings. The approach reduces the number of generator evaluations, improves inference speed by 15.36×, and yields better visual quality, text alignment, and panoramic consistency compared to the strongest training‑free baseline.

By Akio Hayakawa, Yusuke Mukuta, Tatsuya Harada
arXiv Computer Vision
5d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv Computer Vision
Sep 7

Compact Neural Appearance Models for Efficient Gaussian Splatting

The paper introduces a compact neural appearance model for 3D Gaussian Splatting that replaces traditional low‑order spherical harmonics (SH) with a tiny shared MLP decoding per‑primitive latent codes. It compares SH with recent spherical appearance models, integrating all into a unified CUDA rasterizer and WebGL viewer, and demonstrates that the new neural representation reduces per‑primitive appearance storage from 192 to 28 bytes, speeds optimization by 1.3×, and improves reconstruction quality. The study also analyzes how different appearance parametrizations affect geometry recovery and the handling of non‑static scene content.

By Florian Hahlbohm, Jorge Condor, Linus Franke, Martin Eisemann, Marcus Magnor