arXiv:2505.23462v2 Announce Type: replace
Abstract: Blind face restoration from low-quality images is a challenging task that requires not only high-fidelity image reconstruction, but also preservati...
By Runyi Li, Bin Chen, Jian Zhang, Radu Timofte
Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address three key challenges: identity shift, viewpoint-entangled guidance, and perceptual realism.
The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.
By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing
arXiv:2608. 20212v1 Announce Type: new Abstract: High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections.
By Radim Spetlik, David Futschik, Radek Danecek, Feitong Tan, Ziqian Bai, Rohit Pandey, Yinda Zhang
The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.
By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond
arXiv:2603.17825v2 Announce Type: replace
Abstract: In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in vi...
By Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao, Hao Li