CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.01252v1 Announce Type: new Abstract: In camera-controlled video generation, geometry-aware positional encodings condition tokens on camera extrinsics and per-token viewing rays. Existing s...
arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.
We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encod...
The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.
MUGEN is a large‑scale real‑world dataset of over 1,300 hours of 4K panoramic videos with rich semantic and geometric annotations, designed to support interactive 360° world exploration. Wan360 is a camera‑controllable panoramic video generation model built on MUGEN, featuring ERP‑aware components (periodic longitude RoPE, ERP‑aware padding, random roll yaw) and a panoramic Plücker embedding for camera motion. Together, they address gaps in data and model support for immersive, temporally coherent 360° video generation along user‑specified camera trajectories.
The paper introduces CamTrol++, a training‑free method that stabilizes camera‑controlled novel view synthesis from a single image by decomposing large camera motions into small autoregressive steps, thereby limiting per‑step distortion and error accumulation. It also incorporates geometry‑constrained spatial attention, low‑frequency appearance anchoring, and a registration‑free warping pipeline to further enhance stability. Experiments on RealEstate10K and MegaScene demonstrate improved temporal and geometric consistency, better downstream 3D reconstruction quality, and higher generation efficiency, even for long 56‑frame sequences and under depth corruption.