arXiv Machine Learning By Dian Gu, Zhengyi Yang

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

Read the original on arXiv Machine Learning →

arXiv:2606. 07053v1 Announce Type: cross Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.

By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
arXiv Computer Vision
Sep 3

Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters

Make‑It‑Poseable is a feed‑forward framework that treats 3D character posing as a skinning‑free latent‑space transformation. It decouples shape deformation from fixed mesh connectivity, using a latent posing transformer, dense pose representation, and an adaptive completion module with bipartite‑matched latent loss. Experiments show it outperforms existing baselines, generalizes to varied morphologies, and supports 3D authoring tasks such as part replacement and refinement.

By Zhiyang Guo, Ori Zhang, Jax Xiang, Alan Zhao, Zhenxun Yuan, Wengang Zhou, Houqiang Li