arXiv Computer Vision

AstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models

AstraLOD3 is a zero‑shot multimodal agentic system that reconstructs LOD3 building models using multi‑view images, calibrated cameras, a filtered sparse SfM point cloud, and a natural‑language specification. The Astra agent selects and executes computational steps in Python and Blender, achieving a mean FRDS of 0.9647 across 35 runs, including 24 benchmark buildings, with geometric agreement comparable to purpose‑built methods. Ablation studies show the impact of reconstruction guidance, evidence modalities, model configuration, and run‑to‑run variability, demonstrating that LOD3 reconstruction can be framed as a constrained agentic process rather than a fixed pipeline.

arXiv Computer Vision
Sep 23

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

HARMONY is a hierarchical chain-of-thought framework that reconstructs complete 3D indoor scenes from a single monocular image. It combines agentic reasoning with visual geometry foundation models, starting with camera calibration and semantic layout recovery, then placing objects hierarchically while refining geometry with point cloud estimations. The method achieves semantically consistent scenes that align perceptually with the input image, outperforming existing baselines on synthetic and real-world data.

By Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu
arXiv AI
Sep 1

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.

By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
arXiv Computer Vision
Sep 1

SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models

arXiv:2608.29992v1 Announce Type: new Abstract: This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of facade openings in 3D city models, producing directly usable Ci...

By Elmehdi Kanna, Lukas Arzoumanidis, Huynh Duc An Son Nguyen, Youness Dehbi
arXiv AI
Jun 29

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

arXiv:2606. 28215v1 Announce Type: cross Abstract: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.

By Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
arXiv Computer Vision
Sep 22

Semi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms

The paper presents a semi‑automated pipeline that transforms a single 360‑degree video of a classroom into editable, simulation‑ready 3D geometry for CFD analysis. Using NeRF for dense point clouds, SAM‑based 2D instance masks lifted to 3D, and octree‑based instance separation, the workflow produces object‑level assets that can be reconfigured in a browser editor. The resulting geometry is validated with OpenFOAM simulations and benchmarked against IEA Annex 20, demonstrating that per‑room geometry acquisition is essential for accurate airflow predictions.

By Dhruv Gamdha, James Afful, Shambhavi Joshi, Ulrike Passe, Adarsh Krishnamurthy, Baskar Ganapathysubramanian