arXiv:2609.36380v1 Announce Type: new
Abstract: A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program...
By Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
Procedura is a new 3D modeling agent that treats 3D shape as code, using a large language model to generate a procedural assembly from a text prompt. It constructs an assembly graph, writes a parametric program with named parts and typed mates, and verifies each part through compile, mate, and connectivity checks before adding it. A vision critic refines the assembly step‑by‑step, and the resulting program includes per‑part materials and simulator‑validated articulation, producing sharp edges and editable, part‑structured outputs that outperform existing native 3D generators on P3D‑Bench and MechBench‑36.
By Youtian Lin, Yikang Yang, Zhanpeng Hu, Mengqi Zhou, Feihu Zhang, Xun Cao, Jiaheng Liu, Yao Yao
arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.
By Hongyu Wang, Jingquan Wang, Ashvin Anilkumar, Bocheng Zou, Radu Serban, Dan Negrut
PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.
By Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na
arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.
By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
arXiv:2608.21438v1 Announce Type: new
Abstract: Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language d...
By Xiujin Liu, Tianyu Yang, Yilun Zhao, Xiangliang Zhang