Hugging Face Trending Papers

UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step.

arXiv Computer Vision
4d ago

DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization

DiDE is a training‑free framework that enables independent control of color and texture when stylizing 3D assets generated by rectified flow‑based image‑to‑3D models. It leverages the overcomplete latent space of these models, partitioning channels so that texture occupies a small subset while color is encoded in a free subspace. Experiments on the newly collected Disen3D‑Bench benchmark demonstrate that DiDE outperforms existing 2D and 3D stylization methods in color fidelity, texture transfer, and content preservation.

By Tao Wu, Alexandra Gomez-Villa, Senmao Li, Yaxing Wang, Joost van de Weijer, Kai Wang
arXiv Computer Vision
Aug 28

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.

By Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan
arXiv Machine Learning
Jul 16

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

arXiv:2607. 13468v1 Announce Type: cross Abstract: Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation.

By Bin Zang, Wenting Zheng, Xiaoliang Luo, Zhiyuan Fang, Shi Li, Lvchun Wang, Wei Yu, Yi Zhao, Tian Xie, Yuchi Huo, Rengan Xie
arXiv Computer Vision
Aug 31

Cyc3D: Evaluating Cyclic Structural Stability and Asset Usability in Image-to-3D Generation

Cyc3D is a new benchmark for image‑to‑3D generation that evaluates both Cross‑View Object Consistency and Representation Quality. It introduces a closed‑loop View‑Cycle Structural Consistency protocol to measure geometric, perceptual, and semantic drift across repeated render‑regenerate cycles, and also assesses asset usability through geometric structure, fidelity, mesh discretization, and UV quality. Experiments show that closed‑source models outperform open‑source baselines yet still score below 48 on cycle stability, highlighting a gap between visual plausibility and robust 3D understanding.

By Liwen Zhang
arXiv Computation and Language
6d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong