arXiv Computer Vision By Zheng Gao, Xiaoyu Li, Zhicheng Bao, Yang Song, Jiaojiao Jiang

Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering

Read the original on arXiv Computer Vision →

The paper proposes a production‑centered framework for comparing detection and watermarking techniques across two main AI image and video creation routes: direct visual generation and LLM‑driven code rendering. It introduces an explicit verification specification that separates passive inference, message recovery, and authenticated provenance, and organizes watermarks by production stage for images, videos, source code, and rendering‑aware outputs. The authors outline ten research questions covering identifiability, observability, fair comparison, payload recoverability, reconstruction, synchronization, composition, hybrid local contribution, and private production‑event authentication, and connect the framework to concrete systems such as Claude, OpenAI, and rendering tools.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 30

Data Provenance for Image Auto-Regressive Generation

arXiv:2606. 28386v1 Announce Type: cross Abstract: Image autoregressive models (IARs) have recently demonstrated remarkable capabilities in visual content generation, achieving photorealistic quality and rapid synthesis through the next-token prediction paradigm adapted from large language models.

By Bihe Zhao, Louis Kerner, Michel Meintz, Tameem Bakr, Franziska Boenisch, Adam Dziedzic
Hugging Face Trending Papers
Sep 2

Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics

The paper introduces a self-referential retrosynthesis framework for explainable AI provenance forensics that works with fixed generative models. It uses a jointly optimized encoder-decoder pair to embed client inputs, allowing the generator to produce high-fidelity outputs while enabling round-trip consistency checks. The method eliminates the need for watermarking or generator modifications and provides interpretable evidence linking generated images back to their source inputs.

arXiv AI
Sep 2

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

The paper introduces the concept of watermark laundering, where an attacker uses a single reconstruction prompt on public foundation image models to produce a visually faithful output that renders invisible watermarks undecodable. The authors evaluate this failure mode across six OpenAI and Google image editing models, three watermarking schemes, and 1,800 reconstructions, finding that OpenAI models cause the strongest payload disruption while Nano Banana 2 shows vulnerability of DwtDct under high-fidelity reconstruction. Prompt ablation experiments reveal that the disruption is driven by the reconstruction pathway itself rather than any specific removal instruction, highlighting prompt-conditioned reconstruction as a distinct attack interface.

By Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu, Chunpeng Wang, Suo Gao
arXiv AI
Sep 3

Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics

The paper introduces a self‑referential retrosynthesis framework for explainable AI provenance forensics that works with fixed‑generator generative models. It uses a jointly optimized encoder‑decoder pair to embed client inputs, generate high‑fidelity outputs, and then verify provenance by comparing the resynthesized image to the original query. The method eliminates the need for watermarking or generator modifications while providing interpretable evidence of a model’s output origin.

By Yijie Lin, Ching-Chun Chang, Isao Echizen, Hui Li, Chin-Chen Chang
arXiv Computation and Language
Sep 1

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.

By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng
arXiv Computer Vision
Sep 28

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

FeatMark is a watermarking framework that protects images from text‑to‑image diffusion model mimicry attacks by embedding small, scene‑consistent micro‑features instead of pixel‑level perturbations. It constructs domain‑specific feature banks, selects executable features, and injects them via mask‑guided concept editing to create highly localized, natural edits. Experiments on VGGFace2, CelebA‑HQ, and WikiArt show FeatMark remains robust against ten strong watermark removal attacks and several adaptive attacks, with minimal impact on perceptual quality and extending to video mimicry scenarios.

By Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu