arXiv Computer Vision By Bowen Cui, Weijie Wang, Zeyu Zhang, Yefei He, Mingda Lin, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

Read the original on arXiv Computer Vision →

arXiv:2608. 19567v1 Announce Type: new Abstract: While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 3

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

ZipTok3D is a 3D tokenizer that achieves high‑fidelity reconstruction from extremely short token sequences by organizing object geometry into progressively informative global‑token prefixes. During training, nested dropout truncates the latent sequence, forcing each retained prefix to reconstruct the full object, which prioritizes essential geometric information in the leading tokens. The decoder uses a parameter‑shared Transformer block to iteratively recover fine‑grained geometry, enabling reconstruction quality comparable to a 32‑token baseline while using only one token on ShapeNet and four on TRELLIS.

By Mingda Lin, Weijie Wang, Zeyu Zhang, Bowen Cui, Yefei He, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang