arXiv Computer Vision By Bryan G. Pantoja-Rosero

AstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models

Read the original on arXiv Computer Vision →

AstraLOD3 is a zero‑shot multimodal agentic system that reconstructs LOD3 building models using multi‑view images, calibrated cameras, a filtered sparse SfM point cloud, and a natural‑language specification. The Astra agent selects and executes computational steps in Python and Blender, achieving a mean FRDS of 0.9647 across 35 runs, including 24 benchmark buildings, with geometric agreement comparable to purpose‑built methods. Ablation studies show the impact of reconstruction guidance, evidence modalities, model configuration, and run‑to‑run variability, demonstrating that LOD3 reconstruction can be framed as a constrained agentic process rather than a fixed pipeline.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 23

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

HARMONY is a hierarchical chain-of-thought framework that reconstructs complete 3D indoor scenes from a single monocular image. It combines agentic reasoning with visual geometry foundation models, starting with camera calibration and semantic layout recovery, then placing objects hierarchically while refining geometry with point cloud estimations. The method achieves semantically consistent scenes that align perceptually with the input image, outperforming existing baselines on synthetic and real-world data.

By Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu
arXiv AI
Sep 1

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.

By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel