arXiv Machine Learning

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

arXiv:2606. 07053v1 Announce Type: cross Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios.

arXiv AI
Aug 11

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.

By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
arXiv Computer Vision
Sep 3

Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters

Make‑It‑Poseable is a feed‑forward framework that treats 3D character posing as a skinning‑free latent‑space transformation. It decouples shape deformation from fixed mesh connectivity, using a latent posing transformer, dense pose representation, and an adaptive completion module with bipartite‑matched latent loss. Experiments show it outperforms existing baselines, generalizes to varied morphologies, and supports 3D authoring tasks such as part replacement and refinement.

By Zhiyang Guo, Ori Zhang, Jax Xiang, Alan Zhao, Zhenxun Yuan, Wengang Zhou, Houqiang Li
Hugging Face Trending Papers
Sep 8

ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation

ReMoMask-2 is a retrieval‑augmented text‑to‑motion generation framework that improves on complex motion descriptions by addressing coarse retrieval and representation gaps. It introduces a structure‑aware RAG pipeline with Hierarchical Bidirectional Momentum contrastive learning, Semantic Spatial‑Temporal Attention, and Topology Structured Masking, and rebuilds the retrieval database in the generator’s latent space using a lightweight projector. Experiments on HumanML3D, KIT‑ML, and SnapMoGen show state‑of‑the‑art retrieval accuracy and the lowest FID scores, with a single mask‑transformer stage delivering faster inference than the previous two‑stage design.

arXiv Computer Vision
6d ago

Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation

The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.

By Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
arXiv Computer Vision
Aug 25

SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space

SketchFlow is a new generative framework for creating high‑quality vector sketches from text prompts. It uses a Gaussian Mixture Model prior in the CLIP latent space and an Optimal Transport Conditional Flow Matching model to map this prior to sketch features, which are then decoded by a Hybrid Diffusion Decoder combining 1D U‑Net and Transformer architectures. The approach achieves superior visual quality and human‑like drawing styles, and supports zero‑shot synthesis for unseen concepts and smooth semantic interpolation.

By Jin Zhou, Hongliang Yang, Pengfei Xu, Hui Huang