arXiv AI

Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

arXiv:2608. 14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.

arXiv AI
3d ago

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

ReGain is a training‑free correction that improves subject fidelity in text‑to‑image diffusion models personalized with synthetic images. The authors show that fine‑tuning on synthetic images degrades fidelity due to inflated classifier‑free guidance, especially at high frequencies. ReGain measures this inflation per frequency band and scales it down during sampling, closing 51‑64% of the fidelity gap on Stable Diffusion v1.5 and improving performance on SDXL and SD 3.5 while preserving text alignment.

By Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja
arXiv Computer Vision
Sep 11

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.

By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
arXiv Machine Learning
Jun 9

Mitigating Diffusion Model Hallucinations with Dynamic Guidance

arXiv:2510. 05356v2 Announce Type: replace-cross Abstract: Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score function, which in turn leads to interpolations between modes of the data distribution.

By Kostas Triaridis, Alexandros Graikos, Aggelina Chatziagapi, Grigorios G. Chrysos, Dimitris Samaras
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
Hugging Face Trending Papers
Jun 19

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data through data augmentation or image generation.

Hugging Face Trending Papers
Jul 6

LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models

Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each subject's identity while keeping the scene spatially and visually coherent. Methods that fuse independently trained concept adapters in a shared weight space (via federated averaging, gradient fusion, or orthogonality constraints) suffer from identity confusion and style bleeding and require joint retraining.