arXiv AI By Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye

CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Jul 8

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

arXiv:2601. 15275v3 Announce Type: replace-cross Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene.

By Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani
arXiv Computer Vision
Sep 15

G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.

By Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang
arXiv Computer Vision
1d ago

MUGEN: Interactive Panoramic World Exploration via Camera Control

MUGEN is a large‑scale real‑world dataset of over 1,300 hours of 4K panoramic videos with rich semantic and geometric annotations, designed to support interactive 360° world exploration. Wan360 is a camera‑controllable panoramic video generation model built on MUGEN, featuring ERP‑aware components (periodic longitude RoPE, ERP‑aware padding, random roll yaw) and a panoramic Plücker embedding for camera motion. Together, they address gaps in data and model support for immersive, temporally coherent 360° video generation along user‑specified camera trajectories.

By Jiaming Tan, Zhen Li, Shuwei Shi, Minggui Teng, Siqi Yang, Yuwei Wu, Bo Zheng, Chuanhao Li, Kaipeng Zhang
arXiv Computer Vision
Sep 4

Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

The paper introduces CamTrol++, a training‑free method that stabilizes camera‑controlled novel view synthesis from a single image by decomposing large camera motions into small autoregressive steps, thereby limiting per‑step distortion and error accumulation. It also incorporates geometry‑constrained spatial attention, low‑frequency appearance anchoring, and a registration‑free warping pipeline to further enhance stability. Experiments on RealEstate10K and MegaScene demonstrate improved temporal and geometric consistency, better downstream 3D reconstruction quality, and higher generation efficiency, even for long 56‑frame sequences and under depth corruption.

By Prajwal Singh, Arjun Badola, Seema Kumari, Hajime Nagahara, Shanmuganathan Raman