arXiv Computer Vision

Cyc3D: Evaluating Cyclic Structural Stability and Asset Usability in Image-to-3D Generation

Cyc3D is a new benchmark for image‑to‑3D generation that evaluates both Cross‑View Object Consistency and Representation Quality. It introduces a closed‑loop View‑Cycle Structural Consistency protocol to measure geometric, perceptual, and semantic drift across repeated render‑regenerate cycles, and also assesses asset usability through geometric structure, fidelity, mesh discretization, and UV quality. Experiments show that closed‑source models outperform open‑source baselines yet still score below 48 on cycle stability, highlighting a gap between visual plausibility and robust 3D understanding.

arXiv Computer Vision
2d ago

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding introduces SPAR, a joint semantic‑geometric encoding architecture that isolates transient dynamic noise before latent space aggregation. The method couples motion estimation with multi‑view visual and semantic learning in a dynamic‑region‑aware end‑to‑end training paradigm, enabling the network to resolve motion conflicts and produce temporally stable scene representations. Experiments on the D‑RE10K benchmark show state‑of‑the‑art performance, achieving high PSNR values for novel view synthesis and an 88.5% mIoU for motion mask prediction in a self‑supervised setting.

By Boyu Cai, Li Yang, Yan Xu, Wei Liu, Nian Liu, Sikui Zhang, Yan Wang, Chunfeng Yuan, Weiming Hu