arXiv AI

Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding

arXiv:2606. 10892v1 Announce Type: cross Abstract: To showcase products, merchants often incur substantial costs creating high-quality display images.

Hugging Face Trending Papers
Aug 17

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously.

arXiv Computer Vision
6d ago

HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models

HyperErase introduces a hypernetwork-based framework for concept erasure in text-to-image models, replacing static adapters with prompt-conditioned parameter synthesis. The method maps textual descriptions to LoRA updates, eliminating per-prompt gradient optimization and manual merging. A decoupled rectification strategy further stabilizes and refines the synthesized adapters, yielding improved erasure effectiveness, image quality, and semantic alignment across diverse concepts.

By Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
arXiv Computer Vision
Aug 28

Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

Sidecar is a plug‑and‑play semantic augmentation module that preserves entity‑level information from a character’s initial description and injects missing semantics into later prompt embeddings, enabling training‑free character consistency in free‑form visual storytelling. It works with existing diffusion models without architectural changes and incurs negligible computational overhead. Experiments on FreeStoryBench demonstrate consistent improvements in prompt‑image alignment and character consistency across SDXL and FLUX baselines.

By Sibo Dong, Sarah Adel Bargal
arXiv Computer Vision
Sep 11

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.

By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen