arXiv Computer Vision By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

Read the original on arXiv Computer Vision →

The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 7

FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, existing methods still struggle to balance spatial fidelity and temporal coherence under complex degradations. To address this, we propose FADRA, a frequency-aware diffusion framework with iterative residual adaptation specifically tailored for robust VFR.