FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model.
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
arXiv:2602.19202v3 Announce Type: replace Abstract: Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensit...
Forge4D is a feed‑forward model that reconstructs temporally aligned 4D human representations from uncalibrated sparse‑view videos, enabling both novel view and novel time synthesis. It achieves this by jointly streaming 3D Gaussian reconstruction with dense motion prediction, using learnable state tokens for temporal consistency and a self‑supervised retargeting loss for motion prediction. Extensive experiments confirm its effectiveness on in‑domain and out‑of‑domain datasets.
arXiv:2608.22861v1 Announce Type: new Abstract: State Space Models (SSMs) have surfaced as a promising architecture in Video Frame Interpolation (VFI), as they can capture long-range dependencies wit...
arXiv:2609.36940v1 Announce Type: new Abstract: Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are esse...
arXiv:2607. 20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction.
MotionSpec introduces a motion supervision framework for text-to-video generation that focuses on Spectral Trajectory Consistency (STC). STC builds dense anchor-relative motion trajectories, transforms them into spectral volumes, and aligns their amplitude and phase with target trajectories to constrain motion strength and temporal organization. The framework also adds Local Flow Consistency (LFC) to stabilize local motion transitions, resulting in improved motion consistency, temporal coherence, and plausibility while maintaining visual fidelity.
arXiv:2608.20770v1 Announce Type: new Abstract: Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does...
CamWorldQA introduces the first benchmark for assessing the perceptual quality of camera‑controlled world video generation, featuring 720 videos generated by six methods from 20 source videos across six camera trajectories, each scored by human raters. The paper also presents CWQA, a no‑reference quality assessment network that combines spatial, temporal motion, and optical flow features to predict quality scores. Experiments show CWQA outperforms existing VQA methods on the CamWorldQA dataset.
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
arXiv:2609.39504v1 Announce Type: cross Abstract: We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow...