arXiv AI By Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo

Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion

Read the original on arXiv AI →

arXiv:2603. 06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 15

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.

arXiv AI
Sep 3

Bernini: Latent Semantic Planning for Video Diffusion

Bernini proposes a unified framework that separates semantic planning and pixel rendering for video generation and editing. An MLLM-based planner predicts target semantics in ViT embedding space, while a DiT-based renderer synthesizes pixels conditioned on this plan, text features, and source VAE features for editing. The approach introduces Segment-Aware 3D Rotary Positional Embedding and chain-of-thought reasoning, achieving state‑of‑the‑art performance on diverse video benchmarks.

By Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
arXiv Computer Vision
2d ago

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.

By Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang
arXiv Computer Vision
Sep 2

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.

By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo