arXiv AI By Wen Huang, Jiarui Yang, Tao Dai, Jiawei Li, Shaoxiong Zhan, Bin Wang, Shu-Tao Xia

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization

Read the original on arXiv AI →

arXiv:2508. 09459v3 Announce Type: replace-cross Abstract: Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv AI
Sep 7

What Moves? Localized Motion Representations for Compositional Scene Control

The paper introduces a promptable localized motion representation that generates persistent embeddings for user-specified regions in a video, without cropping or masking the input. By conditioning motion encoding directly on spatial masks while processing the full video, the method produces temporally consistent, region-addressable embeddings that capture local dynamics while preserving global context. These embeddings enable object-level motion transfer for dynamic scene composition and improve localized action classification in multi-actor videos, outperforming global representations that rely on cropping or post-hoc masking.

By Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Bj\"orn Ommer