arXiv Computer Vision

Robust, Estimator-Agnostic Dynamic 3DGS Compression

arXiv Computer Vision
Aug 28

KISS-GS: 3D Gaussian Splatting Compression Kept Simple

KISS-GS is a modular compression pipeline for 3D Gaussian Splatting (3DGS) scenes that separates compression from training. It first compacts a vanilla 3DGS scene by 15.7× using state‑of‑the‑art pruning, then encodes the result into the SOG‑XT image‑based format, achieving an additional 6.6× reduction. Optional encoding‑aware fine‑tuning can further cut the size by 2.2×, yielding total reductions of 85× to 319× on standard benchmarks while enabling web‑native decoding.

By Wieland Morgenstern, Friedrich Elias Branschke, Florian Fleischmann, Adrian Szatmari, Paul Schlack, Florian Barthel, Peter Eisert, Anna Hilsmann
arXiv Computer Vision
Sep 16

Racing in Volume with Flow Ensembles

The paper introduces FastFlowGS, a streaming 4D Gaussian Splatting method that reconstructs fast-moving subjects from a sparse set of external cameras, and Monaco4D, a photorealistic Unreal Engine 5 benchmark featuring Formula 1 sequences with dense ground truth. FastFlowGS combines sparse matches, semi-dense tracks, and dense optical flow using a Kalman-style temporal update, achieving significant performance gains over existing baselines on both CMU-Panoptic and Monaco4D datasets. The benchmark provides varied illumination and viewpoints from trackside, onboard, and drone cameras, enabling evaluation of high-speed outdoor reconstruction.

By Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle, Albert Mosella-Montoro, Jose Ribeiro-Gomes, Francisco Vicente Carrasco, Fernando De la Torre
arXiv AI
Aug 21

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

arXiv:2608. 19556v1 Announce Type: cross Abstract: Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion.

By Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
Hugging Face Trending Papers
Aug 3

StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting

Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally.

arXiv Computer Vision
Sep 25

Only What Was Seen: Observation-Gram Compaction of View-Dependent Appearance in 3D Gaussian Splatting

The paper introduces an observation‑Gram matrix that captures how each Gaussian in a 3D Gaussian Splatting model is viewed from training camera directions. This matrix serves as a distortion metric, enabling closed‑form degree reduction, Lagrangian rate‑distortion degree allocation, and matrix‑weighted vector quantisation. When applied to the Compressed3D framework, the metric improves PSNR by 0.49 dB before fine‑tuning and still yields a 0.32 dB gain at matched bitrate without any training images, while a training‑free stack built on the metric is 15% smaller than the image‑free GSICO at equal quality on Mip‑NeRF 360.

By Krzysztof Pietroszek
arXiv Computer Vision
Aug 27

Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-view Videos

Forge4D is a feed‑forward model that reconstructs temporally aligned 4D human representations from uncalibrated sparse‑view videos, enabling both novel view and novel time synthesis. It achieves this by jointly streaming 3D Gaussian reconstruction with dense motion prediction, using learnable state tokens for temporal consistency and a self‑supervised retargeting loss for motion prediction. Extensive experiments confirm its effectiveness on in‑domain and out‑of‑domain datasets.

By Yingdong Hu, Yisheng He, Jinnan Chen, Weihao Yuan, Kejie Qiu, Zehong Lin, Siyu Zhu, Zilong Dong, Steven Hoi, Jun Zhang
arXiv Machine Learning
Jul 1

Drop-In Perceptual Optimization for 3D Gaussian Splatting

arXiv:2603. 23297v2 Announce Type: replace-cross Abstract: Despite their output being ultimately consumed by human viewers, 3D Gaussian Splatting (3DGS) methods often rely on ad-hoc combinations of pixel-level losses, resulting in blurry renderings.

By Ezgi Ozyilkan, Zhiqi Chen, Oren Rippel, Jona Ball\'e, Kedar Tatwawadi
arXiv Computer Vision
4d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai