Hugging Face Trending Papers

InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation

Read the original on Hugging Face Trending Papers →

Text-conditioned human interaction generation must capture both long-range temporal causality within each individual and tightly coupled coordination between partners. Existing interaction diffusion models typically denoise full sequences using bidirectional attention, which obscures causality and hinders streaming and long-horizon generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
5d ago

BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.

arXiv AI
4d ago

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

The paper introduces BRAID, a hierarchical latent-variable model that generates multi-person human motion by explicitly modeling both group-level interaction dynamics and individual behavior conditioned on evolving group context. It treats social motion generation as a meta-transfer learning problem, learning shared interaction priors across datasets and adapting them to arbitrary context sets of observed people and joints. BRAID supports coherent generation under full, sparse, or partial observations and produces compact social-state vectors useful for downstream embodied-agent systems, with evaluations on social forecasting, tracking, in-filling, and response generation.

By Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman
arXiv AI
Jul 29

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

arXiv:2607. 25157v1 Announce Type: new Abstract: Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretraining with bidirectional denoising.

By Zhengtao Yao, Runhao Li, Xupeng Chen, Jiayi Cheng, Chenqian Le, Michael Yue, Jesson Wang, Siheng Wang, Guang Yang, Haoyan Xu, Chenhao Wei, Zhengqing Yuan, Youran Shen, Yanfang Ye, Junhao Dong
arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
Hugging Face Trending Papers
Aug 27

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. By confining bidirectional spatio‑temporal modeling to a fixed‑size window and using bounded temporal and global appearance memories, it emits clean video chunks with low latency. A progressive distillation pipeline further refines the model, achieving superior quality with 26× lower latency and 11× higher throughput compared to prior methods.