arXiv Computer Vision

Delving into Asymmetric Information Dynamics for High-Fidelity Virtual Try-On

The paper introduces RealFit, a virtual try‑on framework that addresses texture degradation and structural drift caused by symmetric bidirectional attention in Diffusion Transformers. By employing Unidirectional Information Flow to isolate garment conditions from stochastic noise and Decoupled Timestep Modulation to strengthen the conditional signal, RealFit achieves high‑fidelity garment rendering. The method also enables a time‑invariant condition branch and a conditional KV cache, cutting inference time by about 75%.

arXiv AI
Jun 11

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

arXiv:2601. 22725v4 Announce Type: replace-cross Abstract: Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck.

By Jin Li, Tao Chen, Kai Wen, Siqi Yin, Shuai Jiang, Weijie Wang, Jingwen Luo, Chenhui Wu
arXiv Computer Vision
Sep 18

Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.

By Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park
arXiv AI
6d ago

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

ReGain is a training‑free correction that improves subject fidelity in text‑to‑image diffusion models personalized with synthetic images. The authors show that fine‑tuning on synthetic images degrades fidelity due to inflated classifier‑free guidance, especially at high frequencies. ReGain measures this inflation per frequency band and scales it down during sampling, closing 51‑64% of the fidelity gap on Stable Diffusion v1.5 and improving performance on SDXL and SD 3.5 while preserving text alignment.

By Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja
arXiv AI
6d ago

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

FoCLIP is a framework that creates adversarial examples to manipulate CLIP-based image quality metrics by reducing the alignment between image and text features. It uses stochastic gradient descent to combine feature alignment, score distribution balancing, and pixel‑guard regularization, enabling high CLIPscore predictions while maintaining visual fidelity. Experiments on artistic prompts and ImageNet show significant CLIPscore gains, and the authors also propose a color‑channel sensitivity detection method that achieves 91% accuracy.

By Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai