arXiv AI

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

The paper introduces MDiTFace, a diffusion transformer designed for high‑fidelity mask‑text collaborative facial generation. It unifies tokenization of semantic masks and text, enabling synchronous multimodal feature interaction via stacked multivariate transformer blocks. A novel decoupled attention mechanism separates dynamic and static computations, allowing caching of static features and reducing mask‑condition overhead by over 94% while preserving performance, leading to superior facial fidelity and conditional consistency compared to existing methods.

arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv AI
Jul 22

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

arXiv:2607. 19344v1 Announce Type: cross Abstract: Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone.

By Rahul Sajnani, Yulia Gryaditskaya, Radom\'ir M\v{e}ch, Srinath Sridhar, Matheus Gadelha
Hugging Face Trending Papers
Jul 23

GroupVideo: Multi-Identity Customized Text-to-Video Generation

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.

arXiv Computer Vision
Sep 3

Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding

The paper introduces MiRA, a plug‑in framework that reweights framewise attention in Vision Transformer video models to better capture subtle facial dynamics for expression recognition. MiRA computes frame‑level confidence and intra‑frame concentration from self‑attention maps, redistributing attention toward localized facial cues without adding trainable parameters. Two modes—an exact post‑softmax redistribution and a lightweight flashLite pre‑softmax approximation—are proposed, and experiments on facial expression recognition benchmarks show consistent gains over strong ViT baselines.

By Seongro Yoon, Donghyeon Cho, Jinsun Park, Fran\c{c}ois Br\'emond
arXiv Computer Vision
Sep 25

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.

By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing
arXiv AI
Sep 12

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

The paper introduces CoMA-DiT, a bidirectional cross‑modal Diffusion Transformer that uses paired modalities as mutual generative supervision for latent augmentation rather than just inputs for fusion. By conditioning velocity prediction on the paired modality through cross‑modal attention and injecting variation via a reliability‑gated residual mechanism, CoMA‑DiT improves multimodal brain state decoding. Experiments on auditory attention decoding and emotion recognition show consistent gains over 20 baselines, with absolute accuracy and macro‑F1 improvements of 4.28% and 6.70% respectively, and extensive analyses confirm its robustness and interpretability.

By Ziwei Wang, Xingyi He, Hongbin Wang, Tianwang Jia, Bohan Fang, Dongrui Wu
arXiv AI
Sep 7

Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

The paper introduces a multimodal emotion recognition framework that combines audio and visual feature extraction with an attention-based fusion strategy. Audio features include Wav2Vec2 embeddings, MFCCs, and statistical acoustic descriptors, fused via a BiLSTM, while video features are extracted using a ResNet50-BiLSTM architecture. A multi-head attention mechanism fuses these modalities, and experiments on MELD and IEMOCAP show significant accuracy and robustness gains, especially in unbalanced data settings.

By Xu Lin, Ke Wang, Hui Kang, Xinying Wang
arXiv Computation and Language
Aug 28

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.

By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim