arXiv Computer Vision By Dhruv Gamdha, James Afful, Shambhavi Joshi, Ulrike Passe, Adarsh Krishnamurthy, Baskar Ganapathysubramanian

Semi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms

Read the original on arXiv Computer Vision →

The paper presents a semi‑automated pipeline that transforms a single 360‑degree video of a classroom into editable, simulation‑ready 3D geometry for CFD analysis. Using NeRF for dense point clouds, SAM‑based 2D instance masks lifted to 3D, and octree‑based instance separation, the workflow produces object‑level assets that can be reconfigured in a browser editor. The resulting geometry is validated with OpenFOAM simulations and benchmarked against IEA Annex 20, demonstrating that per‑room geometry acquisition is essential for accurate airflow predictions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 28

Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras

EgoSurg is a framework that reconstructs dynamic operating room scenes from sparse wall‑mounted stereo video and renders arbitrary, role‑specific egocentric views without instrumenting personnel. It builds a per‑timestamp 3D Gaussian Splatting representation using scale‑aware stereo depth and refines it with an image‑conditioned diffusion model to correct artifacts from limited camera coverage, crowding, and occlusion. Evaluations on real robotic pulmonology procedures and simulated sessions show consistent near‑field reconstruction fidelity (PSNR 26.8 dB, SSIM 0.895) and synthesized egocentric view quality (PSNR 17.8 dB, SSIM 0.766) across workflow phases and hospital sites, with case studies demonstrating applications such as sterile field violation adjudication, role‑specific replay, and counterfactual personnel positioning.

By Han Zhang, Lalithkumar Seenivasan, Jose L. Porras, Roger D. Soberanis-Mukul, Hao Ding, Hongchao Shu, Benjamin D. Killeen, Ankita Ghosh, Lonny Yarmus, Jeffrey K. Jopling, Masaru Ishii, Angela C. Argento, Mathias Unberath
arXiv Computation and Language
Sep 1

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.

By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
arXiv Computer Vision
2d ago

MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation

MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.

By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang