arXiv AI

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

arXiv:2606. 24206v1 Announce Type: cross Abstract: Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model.

arXiv AI
Jul 15

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

arXiv:2607. 12752v1 Announce Type: cross Abstract: While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry.

By Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
arXiv Computer Vision
Sep 7

WorldSculpt: Generating Compositional Worlds from Grounded Videos

WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.

By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
arXiv Computer Vision
Sep 24

Fusion-Aware Direct 3D Gaussian Generation with Structured Patch Latent Flows

The paper introduces a fusion‑aware hierarchical Gaussian patch representation that enables direct class‑guided generation of 3D Gaussian Splatting (3DGS) objects. By decomposing irregular Gaussian sets into canonical local patches and encoding them as structured tokens, the method fuses global class semantics with patch‑level geometry, appearance, spatial correspondence, and rendering‑sensitive cues. A structure‑aware rectified flow model, conditioned on patch positions and coupled with global‑local velocity prediction and density‑aware weighting, produces class‑conditioned 3DGS objects within seconds, achieving more coherent geometry, sharper local details, and better multi‑view consistency than baseline models.

By Yizhao Wang, Jingbo Wang, Guantao Zhang
arXiv Computer Vision
Sep 23

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Fysiverse-3D-Vision is a unified vision‑language‑geometry framework that reconstructs executable 3D scenes from a single image. It separates spatial layout reasoning from asset synthesis, using a shared representation where spatial reasoning and geometric reconstruction reinforce each other. The model employs a Transformer that integrates textual supervision, semantic visual cues, and geometric representations, and includes an object‑conditioned layout module to predict object translation, rotation, and scale while maintaining physical consistency through collision‑aware optimization.

By Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang