arXiv Computer Vision

Training-Free Inpainting Across Domains with a Frozen Text-to-Image Diffusion Model

The paper demonstrates that a frozen generic text‑to‑image diffusion model can perform conditional inpainting across three natural‑image domains without any inpainting‑specific training or dataset adaptation. The proposed Step‑PI method augments known‑region projection with boundary‑interior latent feedback, persistent PI state, and a predefined four‑field release schedule, improving all evaluated metrics over baseline training‑free approaches. Experiments on CelebA‑HQ, AFHQ, and Places2 show consistent gains, with Step‑PI outperforming LanPaint and PILOT on all five macro metrics.

arXiv AI
3d ago

Image AID via continuous-time reinforcement learning

The paper introduces Amortized Inpainting with Diffusion (AID), a method that keeps a pretrained diffusion backbone fixed and trains a small reusable guidance module offline for image inpainting. AID formulates the problem as deterministic guidance with a supervised terminal objective, derives an auxiliary Gaussian formulation to make it learnable, and proves that solving the randomized problem recovers the optimal deterministic guidance field. Experiments on AFHQv2, FFHQ, and ImageNet show that AID improves the quality–speed trade‑off over strong baselines while adding less than one percent trainable overhead.

By Yilie Huang, Xun Yu Zhou
arXiv AI
Sep 15

Freeze, Share, Shrink: Rethinking the Action Backbone in Diffusion Policies

The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.

By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
arXiv AI
3d ago

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

CAST introduces a reinforcement‑learning fine‑tuning framework for diffusion models that addresses three key limitations: it automatically selects the denoising window based on each model’s trajectory, decomposes prompts into verifiable semantic atoms via Causal Scene Graphs, and applies atom‑level rewards spatially weighted in the policy objective. The method is applied to FLUX.2‑dev and Qwen‑Image‑2512, yielding up to 3.07× improvement on the hardest GenEval 2 prompts compared with Flow‑GRPO while also enhancing overall generation quality.

By Shu Yu, Chaochao Lu
arXiv Computer Vision
Aug 26

On-Policy Self-Distillation in Diffusion Models

The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.

By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
arXiv Computer Vision
Sep 11

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.

By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
Hugging Face Trending Papers
Jul 6

RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement

Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object pairings, often resulting in concept omission or semantic drift where a dominant entity overwhelms the generation. Tracing these failures to a lack of compositional balance during the denoising trajectory, we propose RADIANCE, a training-free framework that treats inference as a closed-loop feedback process.

arXiv Computer Vision
Sep 3

Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

The paper introduces a structured‑prior‑guided diffusion inpainting framework that enhances traffic sign augmentation by incorporating semantic, appearance, and geometric priors through text prompts, vector templates, and ControlNet. It enforces physical consistency with colour and edge‑structure losses, achieving superior reconstruction fidelity and OCR accuracy compared to larger models. The synthetic data generated improves detection performance for rare traffic sign classes by up to 7.4× over a real‑data‑only baseline.

By Luo Li, Chongchong Huang, Jun Jia, Qiang Gao, Xinlong Liu, Gui Yang, Liang Cao