arXiv Machine Learning

Amortized Moment Matching for Visual Generation

arXiv:2607. 26860v1 Announce Type: new Abstract: We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals.

arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried
Hugging Face Trending Papers
Aug 6

Energy-Guided Flow Matching

Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly.

arXiv Computer Vision
Sep 14

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion. "whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."

By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
arXiv Computer Vision
Sep 17

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

FlashAR is a lightweight post‑training adaptation framework that converts a pre‑trained raster‑scan autoregressive image model into a highly parallel generator using two‑way next‑token prediction. It preserves the original training objective by keeping the horizontal head for row‑wise prediction and adding a lightweight vertical head for column‑wise prediction, with a learnable fusion gate to combine the two predictions. A two‑stage adaptation pipeline—first initializing the vertical head from the pre‑trained model and then jointly fine‑tuning—yields up to a 22.9× speedup for 512×512 image generation while using only 0.05% of the original training data.

By Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang