arXiv Computer Vision By Junyu Zhu, Hao Zhu, Xinzhuo Zhang, Xu Zhang, Hongdong Li, Zhan Ma, Xun Cao

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

Read the original on arXiv Computer Vision →

ASTRA (Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment) tackles the challenge of reconstructing dynamic 3D scenes from temporally asynchronous multi‑camera data. By using 2D motion trajectories as texture‑robust supervision, it jointly optimizes temporal offsets and 3D representations, aligning projected 3D point motion with observed 2D paths while masking unreliable constraints. Experiments on Gaussian Splatting backbones show that ASTRA retains high‑frequency spatial detail, improves PSNR by ~1.4 dB, reduces temporal‑offset MAE by 54 %, and nearly quadruples synchronization success even with up to 25‑frame offsets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Aug 11

CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images

Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and motion blur. Although neural rendering has advanced sparse-view synthesis, existing blur-aware methods typically require substantial multi-view redundancy, accurate camera poses, or costly per-scene optimization.

arXiv Computer Vision
Aug 24

Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving

The paper introduces Driving with DINO (DwD), a framework that uses Vision Foundation Module (VFM) features to bridge simulation and real-world domains for autonomous driving video generation. It addresses the consistency‑realism dilemma by projecting VFM features onto a principal subspace, dropping high‑frequency texture elements, and applying a Random Channel Tail Drop to preserve structural detail. Additional components— a learnable Spatial Alignment Module and a Causal Temporal Aggregator— enhance control precision, spatial alignment, and temporal stability, reducing motion blur and ensuring realistic, consistent outputs.

By Xuyang Chen, Conglang Zhang, Chuanheng Fu, Zihao Yang, Kaixuan Zhou, Yizhi Zhang, Yanfeng Zhang, Mingwei Sun, Zhen Dong, Xiaoxiao Long, Zengmao Wang, Liqiu Meng
arXiv AI
Jun 2

A Survey of 3D Reconstruction with Event Cameras

arXiv:2505. 08438v4 Announce Type: replace-cross Abstract: Event cameras are rapidly emerging as powerful vision sensors for 3D reconstruction, uniquely capable of asynchronously capturing per-pixel brightness changes.

By Chuanzhi Xu, Haoxian Zhou, Langyi Chen, Haodong Chen, Zeke Zexi Hu, Zhicheng Lu, Ying Zhou, Vera Chung, Qiang Qu, Weidong Cai