Hugging Face Trending Papers

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating high‑quality, globally coherent texture maps for 3D garment reconstruction from a single image. It first generates a coarse texture initialization on the garment’s sewing pattern using a 3D mesh and Vision‑Language Model priors, then refines this in the UV domain with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments that outperform existing baselines.

arXiv Computer Vision
Sep 25

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating production‑ready 3D garment assets from a single image. It generates globally coherent texture maps directly in the 2D sewing pattern (UV) space, using a coarse initialization from Vision‑Language Models and refining it with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments with high‑quality textures that outperform current state‑of‑the‑art baselines.

By Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang, Hugo Bertiche, Alexandru-Eugen Ichim, Thabo Beeler, Fernando De la Torre
arXiv AI
Sep 1

GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns

GarmentWeaver is a new framework for multimodal sewing pattern generation that uses a schema‑aware approach to construct compact hierarchical targets. By activating garment‑relevant structural branches and building on a pretrained vision‑language model, it predicts executable sewing patterns in a structured manner. Experiments show that GarmentWeaver produces more accurate, executable patterns and yields better simulation results than strong baselines.

By Yinwen Lu, Weihao Luo, Yueqi Zhong
arXiv Computer Vision
Sep 11

Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need

The paper introduces AvaImg, a multi‑stage optimization pipeline that achieves high‑fidelity SMPL(-X)+D registrations with UV texture for arbitrary clothed scans. By enforcing a body‑inside‑clothing constraint through signed winding numbers and employing a three‑level efficiency cascade, AvaImg significantly reduces runtime and storage while recovering fine surface detail via coarse‑to‑fine displacement optimization. The resulting textured registrations are nearly indistinguishable from scans, and encoding the UV maps with a frozen FLUX VAE demonstrates compatibility with 2D generative models, enabling 3D avatar generation using image‑based priors.

By Margaret Kostyrko, Yuxuan Xue, Garvita Tiwari, Gerard Pons-Moll
arXiv Computer Vision
4d ago

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

DirectUV is an image-conditioned UV texture diffusion framework that generates high-quality textures directly in UV space using a pretrained image VAE and a Diffusion Transformer. It introduces Surface-Aware Positional Encoding (SAPE), which replaces standard 2D-grid positional encoding with encodings derived from 3D surface coordinates, enabling attention to operate based on surface proximity rather than UV-grid distance. A multi-level extension assigns different attention heads to progressively finer subdivisions of the UV patch, allowing the model to reason about surface structure at multiple granularities, resulting in sharper and more globally consistent textures, especially in occluded and view-unseen regions.

By Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao Chen, Ying-Cong Chen
arXiv Computer Vision
Sep 4

PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models

PoseDreamer is a new pipeline that uses diffusion models to generate large‑scale synthetic datasets for 3D human mesh estimation, providing 3D mesh annotations that remain aligned with the generated images. The system incorporates controllable image generation, Direct Preference Optimization for control alignment, curriculum‑based hard sample mining, and multi‑stage quality filtering to produce over 500,000 high‑quality samples with a 76% improvement in image‑quality metrics over traditional rendering‑based datasets. Models trained on PoseDreamer match or surpass those trained on real‑world or conventional synthetic data, and combining PoseDreamer with synthetic datasets yields better performance than mixing real and synthetic data alone.

By Lorenza Prospero, Orest Kupyn, Ostap Viniavskyi, Jo\~ao F. Henriques, Christian Rupprecht