Hugging Face Trending Papers

Detect Early, Escalate Rarely: Anytime Detection of AI-Generated Video from the Compressed Bitstream

Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasingly by a large vision-language model.

arXiv Machine Learning
Sep 25

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.

By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv Computation and Language
Sep 4

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.

By Prakhar Khatri
arXiv AI
Jul 7

SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness

arXiv:2607. 02886v1 Announce Type: cross Abstract: Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsely rejecting authentic content, a regime where standard metrics such as AUROC fail to reflect actual operating behavior.

By Jongyeop Hyun, Hyounghun Kim
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
arXiv Computer Vision
Sep 24

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

The paper introduces the concept of Information Capacity (IC) to quantify how much bandwidth savings a unit of decoder compute can achieve in generative video compression (GVC). By modeling reconstruction quality as a two‑factor power law in data rate and compute, the authors fit measured DISTS of two GVC decoders with high accuracy and define IC as the negative logarithmic slope along an iso‑quality contour. IC is dimensionless, enabling architecture‑agnostic comparisons and revealing that a 14B decoder trades compute for rate far more efficiently than a 1.3B decoder, with significant variation across datasets.

By Cheng Yuan, Jiawei Shao, Xuelong Li
arXiv AI
Sep 15

Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models

The paper introduces the Compute-Value Audit (CVA), a sequential framework that evaluates whether extra sampling during test‑time scaling for video world models actually yields a net benefit after accounting for the compute cost of generation and verification. On 192 Physics‑IQ scenes, increasing the sample pool from 4 to 16 candidates improves oracle quality by +9.23 IQ, yet common metrics such as Flow, Cycle, and VideoReward fail to reliably recover this headroom, and adaptive‑depth policies recover only 42‑69% of the potential gain. Only a few specific interventions—anchor‑explorer in a sparse PRM800K setting, MMLU‑Pro exposing a predictive‑state gap, and a privileged paired‑future upper bound—successfully pass all CVA stages, indicating that sampling headroom is valuable only when it can be converted into a reliable decision that survives the full compute charge.

By Yuhua Jiang, Junjie Lu, Feifei Gao
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv AI
Sep 25

Accelerating Video Diffusion via Training-Free Trajectory Routing

The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.

By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh