arXiv:2610.00600v1 Announce Type: new
Abstract: Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-i...
By Yuyao Zhang, Yuwei Hu, Ziyang Mai, Yu-Wing Tai
Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion.
"whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."
By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
arXiv:2609.36374v1 Announce Type: new
Abstract: Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research gro...
By Brandon Leblanc, Charalambos Poullis
arXiv:2609.36348v1 Announce Type: cross
Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
By Xiaoyu Wu, Yifei Wang, Chen Wei
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.
arXiv:2603.15132v3 Announce Type: replace
Abstract: While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the raw pixel m...
By Hainuo Wang, Mingjia Li, Xiaojie Guo