arXiv:2609.07137v1 Announce Type: cross
Abstract: Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training st...
By Zhiwei Ning, Zhen Zhou, Puhua Jiang, Xintong Han, Gengming Zhang, Jie Yang, Zhonglong Zheng, Yuanjie Zheng, Wei Liu, Chunchao Guo
arXiv:2610.01233v1 Announce Type: new
Abstract: Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Repre...
By Zhen Zhou, Zhiwei Ning, Puhua Jiang, Sheng Zhang, Yifei Tang, Jie Yang, Xintong Han, Wei Liu, Chunchao Guo
The paper introduces the Geometry‑Native Autoencoder (GAE), a compact latent space that can be decoded into appearance, depth, camera parameters, and point maps, enabling 3D‑consistent world generation. By reparameterizing a geometry foundation model’s features, GAE replaces traditional appearance‑centric latents and improves visual quality and 3D coherence, achieving significant reductions in FVD and camera‑trajectory error on benchmark datasets. The work demonstrates that a geometry‑native latent space can serve as a shared interface between perception and generation models.
By Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2607.05568v2 Announce Type: replace-cross
Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
AGORA is a new framework that extends 3D Gaussian Splatting with a generative adversarial network to produce high‑fidelity, animatable 3D head avatars. It introduces a lightweight FLAME‑conditioned deformation branch that predicts per‑Gaussian residuals for identity‑preserving, fine‑grained expression control, and a dual‑discriminator training scheme that enforces expression fidelity. The system achieves real‑time inference at 250 FPS on a single GPU and, for the first time, CPU‑only animatable 3DGS avatar synthesis at ~9 FPS.
By Ramazan Fazylov, Sergey Zagoruyko, Aleksandr Parkin, Stamatis Lefkimmiatis, Ivan Laptev