arXiv:2606.13345v2 Announce Type: replace
Abstract: Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipeli...
By Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong, Jiachen Xu, Yuan Xie, Xin Tan
arXiv:2605.04412v3 Announce Type: replace
Abstract: 3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects fro...
By Yiran Qiao, Yiren Lu, Yunlai Zhou, Disheng Liu, Linlin Hou, Rui Yang, Yu Yin, Jing Ma
arXiv:2608.25461v1 Announce Type: cross
Abstract: Using conditional image generators, texture artists can explore many single-view looks for an existing 3D shape. Despite impressive progress, state-o...
By Chenyue Cai, Anita Hu, James Lucas, Szymon Rusinkiewicz, Masha Shugrina
Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step.
arXiv:2609.23169v1 Announce Type: new
Abstract: High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown prom...
By Yibo Zhang, Ze Yuan, Nan Cao, Li Zhang, Yan-Pei Cao, Yuan-Chen Guo, Rui Ma
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, alpha mattes, masks, and heatmaps are encoded into an RGB-trained VAE latent space and decoded back as image-like targets.
SafeStyle is a training‑free framework that injects calibrated style residuals into frozen diffusion models for reference‑guided stylization. It estimates style‑supported and content‑associated subspaces from small calibration sets, then transports purified style evidence across adaptive spatial granularity while limiting its influence with a residual‑norm budget. Experiments on texture‑ and geometry‑dominant styles show high DINO style similarity (0.432–0.474) with minimal semantic leakage (0.8%).
By Zhangping Yang, Min Li, Song Yan, Rong Gao, Xinliang Bi, Guanye Xiong, Yujie He
arXiv:2609.38136v1 Announce Type: new
Abstract: Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects,...
By Teng Zhou, Yunhao Chen
The paper introduces Chameleon, a two‑stage training framework for cross‑domain image compositing that separates style and content representations. It first trains a ChameleonEncoder using Joint Hard Contrastive Learning to disentangle style and content, then applies Spatio‑Temporal Attention Gating within a diffusion transformer to stylize the foreground while preserving its identity. The authors also release ChameleonDataset, the first large‑scale training set for cross‑domain compositing, and demonstrate that Chameleon outperforms existing in‑domain, cross‑domain, and commercial models in both plausibility and stylistic fidelity.
By Sukhun Ko, Soo Ye Kim, Jihyong Oh
DecomVoxel introduces a guided in‑situ denoising optimization that fuses 3D‑native priors with neural scene reconstruction to improve decompositional scene reconstruction. The method employs an epsilon‑based distillation loss for stable latent refinement and adaptive spatial guidance using occupied and vacant anchors with temporal annealing to reduce hallucinations and spatial drift. Experiments on Replica and ScanNet++ demonstrate that DecomVoxel outperforms state‑of‑the‑art approaches while preserving spatial layout, structural fidelity, and style‑consistent texture, yielding high‑quality textured meshes with clean topology.
By Junfeng Ni, Zirui Zhou, Yixin Chen, Yu Liu, Nan Jiang, Zhifei Yang, Song-Chun Zhu, Siyuan Huang
arXiv:2606. 07117v1 Announce Type: cross Abstract: This paper presents Native3D, the first end-to-end 3D scene generation framework that completely bypasses 2D intermediate representations.
By Yibo Liu, Ziwei Zhang, Haozhou Pang, Menghao Li, Lanshan He, Gan Qi
The paper introduces a framework for instruction‑guided 3D editing that does not require paired 3D supervision. It distills visual, semantic, and geometric knowledge from foundation models into a 3D editing model using a differentiable rendering pipeline, guided by a 2D visual prior from an image editing model and a semantic prior from a Vision‑Language Model. A 3D‑aware Distribution Matching regularization is added to prevent geometric collapse and ensure realistic 3D outputs, leading to superior instruction fidelity and cross‑view consistency compared to state‑of‑the‑art baselines.
By Hao Wen, Weibin Yun, Hongxing Fan, Haotian Lu, Rui Chen, Zehuan Huang, Lu Sheng