arXiv Computer Vision

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

arXiv:2607. 17097v2 Announce Type: replace Abstract: Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI.

arXiv Computer Vision
Aug 27

4DStreamCtrl: Interactive Video Generation with Online 4D Control

The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.

By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
arXiv Computer Vision
3d ago

Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update

arXiv:2610.07031v1 Announce Type: new Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...

By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv Computer Vision
5d ago

MoSE3: Learning World-Space SE(3) at Every Pixel

MoSE3 is a feed‑forward model that predicts dense SE(3) motion—full 6‑DoF rigid transforms—at every pixel from monocular RGB video, providing rotation, translation, and grouping information simultaneously. It achieves this by jointly learning 3D point tracks and rigidity embeddings, then differentiably fitting transforms within soft rigid clusters. The authors also release Art‑Kubric, a large synthetic dataset with dense SE(3) and rigidity labels, and demonstrate state‑of‑the‑art performance on both rigid and articulated benchmarks, with strong generalization to real‑world videos.

By Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang
arXiv Computer Vision
Sep 1

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
arXiv Machine Learning
Sep 14

AnyView: Synthesizing Any Novel View in Dynamic Scenes

AnyView is a diffusion-based video generation framework designed for dynamic view synthesis, requiring minimal inductive biases or geometric assumptions. It trains a generalist spatiotemporal implicit representation using diverse data sources—monocular, multi-view static, and multi-view dynamic—to produce zero-shot novel videos from arbitrary camera locations and trajectories. The authors evaluate AnyView on standard benchmarks, introduce a new challenging benchmark called AnyViewBench for extreme dynamic view synthesis, and demonstrate that AnyView outperforms baselines in maintaining realistic, plausible, and spatiotemporally consistent videos across diverse real-world scenarios.

By Basile Van Hoorick, Dian Chen, Shun Iwase, Pavel Tokmakov, Muhammad Zubair Irshad, Igor Vasiljevic, Swati Gupta, Fangzhou Cheng, Sergey Zakharov, Vitor Campagnolo Guizilini
arXiv Computer Vision
Sep 7

Learning 3D Editing without Paired Supervision via Generative Prior Distillation

The paper introduces a framework for instruction‑guided 3D editing that does not require paired 3D supervision. It distills visual, semantic, and geometric knowledge from foundation models into a 3D editing model using a differentiable rendering pipeline, guided by a 2D visual prior from an image editing model and a semantic prior from a Vision‑Language Model. A 3D‑aware Distribution Matching regularization is added to prevent geometric collapse and ensure realistic 3D outputs, leading to superior instruction fidelity and cross‑view consistency compared to state‑of‑the‑art baselines.

By Hao Wen, Weibin Yun, Hongxing Fan, Haotian Lu, Rui Chen, Zehuan Huang, Lu Sheng