arXiv Computer Vision

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

DirectUV is an image-conditioned UV texture diffusion framework that generates high-quality textures directly in UV space using a pretrained image VAE and a Diffusion Transformer. It introduces Surface-Aware Positional Encoding (SAPE), which replaces standard 2D-grid positional encoding with encodings derived from 3D surface coordinates, enabling attention to operate based on surface proximity rather than UV-grid distance. A multi-level extension assigns different attention heads to progressively finer subdivisions of the UV patch, allowing the model to reason about surface structure at multiple granularities, resulting in sharper and more globally consistent textures, especially in occluded and view-unseen regions.

arXiv Computer Vision
Sep 25

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating production‑ready 3D garment assets from a single image. It generates globally coherent texture maps directly in the 2D sewing pattern (UV) space, using a coarse initialization from Vision‑Language Models and refining it with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments with high‑quality textures that outperform current state‑of‑the‑art baselines.

By Ding-Jiun Huang, Yuanhao Wang, Cheng Zhang, Hugo Bertiche, Alexandru-Eugen Ichim, Thabo Beeler, Fernando De la Torre
Hugging Face Trending Papers
Sep 24

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

OmniFabric is a new method for creating high‑quality, globally coherent texture maps for 3D garment reconstruction from a single image. It first generates a coarse texture initialization on the garment’s sewing pattern using a 3D mesh and Vision‑Language Model priors, then refines this in the UV domain with a diffusion transformer conditioned on 3D positional features. The approach removes distortion and baked‑in artifacts, producing photorealistic 3D garments that outperform existing baselines.

arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv Computer Vision
Sep 11

HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition

HiPerViT is a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based pipeline for texture recognition. It combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. Across six texture recognition benchmarks, HiPerViT consistently outperforms strong vision-only baselines, achieving notable gains on DTD, GTOS-Mobile, and 1200Tex, and the improvements are largely independent of backbone depth or fusion topology.

By Jo\~ao Pedro C. A. de S\'a, Odemir Martinez Bruno
arXiv Computer Vision
Sep 2

JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

JanusMesh introduces a fast, training‑free framework for creating 3D visual illusion meshes that reveal different semantics from various viewpoints. The method splits generation into two stages: a cross‑space dual‑branch denoising process that aligns 3D latents with CLIP guidance and blends Signed Distance Fields for seamless geometry, followed by a view‑conditioned texture synthesis module that aggregates 2D diffusion priors onto the fused mesh. Experiments show that JanusMesh produces highly realistic, dual‑semantic 3D illustrations in only 3–5 minutes, outperforming prior approaches in geometric integrity, semantic recognizability, and efficiency.

By Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang, Yu-Lun Liu