MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.
BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.
The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.
arXiv:2608. 09226v1 Announce Type: cross Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression.
arXiv:2609.36995v1 Announce Type: cross Abstract: Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-relate...
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).