arXiv Computer Vision

TRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection

arXiv AI
Jun 2

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

arXiv:2606. 00101v1 Announce Type: cross Abstract: With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security.

By Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
arXiv Computer Vision
2d ago

RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection

The paper introduces RED (Reconstruction Evolution Dynamics), a new framework for detecting AI-generated images that leverages the evolution of intermediate reconstruction stages rather than relying solely on static representations or endpoint discrepancies. RED uses a frozen multiscale VQ‑VAE and a frozen CLIP encoder to capture a reconstruction trajectory, then learns image‑adaptive stage weights from token negative log‑likelihoods provided by a frozen VAR model. Experiments on six benchmarks show RED achieves the highest average accuracy (92.5%) and precision (97.5%) among evaluated methods, and it remains robust to common image degradations.

By Wenpeng Mu, Junshan Jin, Tanfeng Sun, Xinghao Jiang, Qiang Xu
Hugging Face Trending Papers
Aug 13

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization.

arXiv AI
Aug 28

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.

By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
arXiv Computer Vision
4d ago

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.

By Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem, Aykut Erdem, Duygu Ceylan