arXiv AI By Karthik Mohan Kumar, Damian Andrysiak, Pedro Antonio Pena, Kunal Tyagi, Rama Harihara

DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering

Read the original on arXiv AI →

DAGS introduces a lightweight, attention‑free conditioning scheme that disentangles appearance and geometry for a frozen image diffusion transformer (DiT), enabling high‑fidelity, temporally stable renders with independent control. Two small convolutional encoders generate per‑frame conditioning features, which are injected as learned residuals into the image tokens, avoiding the quadratic cost of attention. Coupled with a recurrent lighting stabilizer and a training‑free temporal guidance term, DAGS transforms a per‑frame image model into a streaming renderer that outperforms real‑time denoisers and diffusion renderers in PSNR and temporal stability while requiring far less compute than path tracing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 7

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.

Hugging Face Trending Papers
Aug 18

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

MoE-ViE introduces a Mixture-of-Experts vision encoder that scales efficiently for image and video understanding, outperforming dense counterparts across various sizes. The study shows fine‑grained MoE topologies provide significant gains, and proposes an auxiliary‑loss‑free balancing variant and a specialized MoE kernel to reduce inference latency. With frame‑level distillation and a novel freezing mechanism, the largest MoE‑ViE model matches state‑of‑the‑art zero‑shot performance while being 1.7× larger and 76% faster, and it outperforms other encoders when paired with a language model on both image and video benchmarks.

arXiv Computer Vision
Sep 22

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

arXiv:2609.23153v1 Announce Type: new Abstract: Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}:...

By Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui
arXiv Computer Vision
Sep 4

SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

The paper introduces SPARK, a lightweight input‑conditioned controller that modulates only a few dominant channels in frozen Diffusion Transformer (DiT) based super‑resolution models. By predicting bounded per‑channel affine transformations for selected channels, SPARK improves reconstruction fidelity and perceptual quality without fine‑tuning the backbone or adding adapters. Experiments on three DiT‑based SR backbones across DIV2K, RealSR, and DRealSR demonstrate consistent gains while modulating only eight channels per stream and block.

By Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi