EgoSurg is a framework that reconstructs dynamic operating room scenes from sparse wall‑mounted stereo video and renders arbitrary, role‑specific egocentric views without instrumenting personnel. It builds a per‑timestamp 3D Gaussian Splatting representation using scale‑aware stereo depth and refines it with an image‑conditioned diffusion model to correct artifacts from limited camera coverage, crowding, and occlusion. Evaluations on real robotic pulmonology procedures and simulated sessions show consistent near‑field reconstruction fidelity (PSNR 26.8 dB, SSIM 0.895) and synthesized egocentric view quality (PSNR 17.8 dB, SSIM 0.766) across workflow phases and hospital sites, with case studies demonstrating applications such as sterile field violation adjudication, role‑specific replay, and counterfactual personnel positioning.
By Han Zhang, Lalithkumar Seenivasan, Jose L. Porras, Roger D. Soberanis-Mukul, Hao Ding, Hongchao Shu, Benjamin D. Killeen, Ankita Ghosh, Lonny Yarmus, Jeffrey K. Jopling, Masaru Ishii, Angela C. Argento, Mathias Unberath
arXiv:2610.01863v1 Announce Type: new
Abstract: We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes f...
By Zhening Huang, Yueyan Li, Johnathan Chiu, Xiaoyang Lyu, Matt Zhou, Yuxin Yao, Joan Lasenby, Shangzhe Wu
arXiv:2608.30821v1 Announce Type: cross
Abstract: Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and e...
By Minghan Qin, Yuang Wang, Xiuyu Yang, Yushi Long, Yujian Zhang, Ruihuan Wang, Kai Ye, Yangang Zhang, Hang Li
PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.
By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.
By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang
arXiv:2605. 10873v2 Announce Type: replace-cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existing evaluations are fragmented across datasets, modalities, and metrics.
By Anna C. Doris, Jacob Thomas Sony, Ghadi Nehme, Era Syla, Amin Heyrani Nobari, Faez Ahmed
arXiv:2608.30423v1 Announce Type: cross
Abstract: Splatting-based algorithms reconstruct photorealistic, real-time-renderable, and mesh-exportable 3D scenes from regular images, but they represent a...
By Minhas Kamal, Hiranya Garbha Kumar, Mahedi Kamal, Balakrishnan Prabhakaran
arXiv:2608.11645v2 Announce Type: replace
Abstract: Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame...
By Hossein Khalili (UCLA), Philip Do (UCLA), Alexander Vilesov (UCLA), Achuta Kadambi (UCLA), Kittipat Apicharttrisorn (Nokia Bell Labs), Nader Sehatbakhsh (UCLA)
AstraLOD3 is a zero‑shot multimodal agentic system that reconstructs LOD3 building models using multi‑view images, calibrated cameras, a filtered sparse SfM point cloud, and a natural‑language specification. The Astra agent selects and executes computational steps in Python and Blender, achieving a mean FRDS of 0.9647 across 35 runs, including 24 benchmark buildings, with geometric agreement comparable to purpose‑built methods. Ablation studies show the impact of reconstruction guidance, evidence modalities, model configuration, and run‑to‑run variability, demonstrating that LOD3 reconstruction can be framed as a constrained agentic process rather than a fixed pipeline.
By Bryan G. Pantoja-Rosero
arXiv:2605. 01171v2 Announce Type: replace-cross Abstract: Despite recent progress, recovering parametric CAD construction sequences from geometric input, such as meshes or point clouds, is a key challenge for design and manufacturing, as existing CAD reconstruction and generation methods are largely restricted to difficult-to-edit formats like meshes or Breps or editable simple sketch-and-extrude pipelines and low-complexity datasets.
By Ghadi Nehme, Eamon Whalen, Faez Ahmed
ChannelFlow-Tools is an open‑source, configuration‑driven pipeline that generates machine‑learning‑ready datasets for three‑dimensional obstructed channel flows. It combines procedural obstacle geometry generation across six shape families, signed‑distance‑field voxelisation, lattice‑Boltzmann simulation, and packaging into ML‑ready tensors, all driven by reproducible configuration files. The pipeline is validated through mesh‑integrity audits, SDF representation checks, solver benchmarks, and data‑integrity audits, and it has been used to train surrogate models (3D U‑Net, FNO, U‑FNO) that learn geometry‑to‑flow mappings and exhibit physically interpretable behaviour on out‑of‑distribution splits.
By Shubham Kavane, Lukas Schr\"oder, Kajol Kulkarni, Fernando Gonzalez, Harald Koestler
Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation.