Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.
arXiv:2608.23549v1 Announce Type: new
Abstract: Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces...
By Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
MoE-ViE introduces a Mixture-of-Experts vision encoder that scales efficiently for image and video understanding, outperforming dense counterparts across various sizes. The study shows fine‑grained MoE topologies provide significant gains, and proposes an auxiliary‑loss‑free balancing variant and a specialized MoE kernel to reduce inference latency. With frame‑level distillation and a novel freezing mechanism, the largest MoE‑ViE model matches state‑of‑the‑art zero‑shot performance while being 1.7× larger and 76% faster, and it outperforms other encoders when paired with a language model on both image and video benchmarks.
arXiv:2609.23153v1 Announce Type: new
Abstract: Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}:...
By Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui
arXiv:2609.23169v1 Announce Type: new
Abstract: High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown prom...
By Yibo Zhang, Ze Yuan, Nan Cao, Li Zhang, Yan-Pei Cao, Yuan-Chen Guo, Rui Ma
The paper introduces SPARK, a lightweight input‑conditioned controller that modulates only a few dominant channels in frozen Diffusion Transformer (DiT) based super‑resolution models. By predicting bounded per‑channel affine transformations for selected channels, SPARK improves reconstruction fidelity and perceptual quality without fine‑tuning the backbone or adding adapters. Experiments on three DiT‑based SR backbones across DIV2K, RealSR, and DRealSR demonstrate consistent gains while modulating only eight channels per stream and block.
By Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi