Hugging Face Trending Papers

Clean-Reference Streaming Detection of Lens Occlusion and Photometric Transitions for Camera Tamper Monitoring

A surveillance camera is an image sensor whose silent physical degradation invalidates every downstream consumer of its data. In-situ integrity alarms for such vision sensors require low false-alarm rates, bounded computation, and diagnosable behavior under nuisance illumination changes.

arXiv Computer Vision
Sep 4

Hold-Out Self-Validation Cannot Certify Photogrammetric Accuracy: Saturation and Blindness to Coherent Distortion

The paper argues that internal self-consistency checks cannot guarantee the accuracy of photogrammetric reconstructions, a limitation that is structural rather than a tuning issue. It introduces a track‑leakage‑free hold‑out protocol that withholds a deterministic subset of images and tests each against only 3D points supported by at least two retained images, ensuring no view is evaluated against the structure it helped create. Experiments on diverse datasets show that while the protocol is well‑posed, it saturates at a confidence score of 1.00 and fails to detect coherent distortion, missing large errors that can reach over 100 m. whyItMatters":"The study highlights that hold‑out self‑validation scores, increasingly used as quality evidence for metric deliverables, may be misleading and cannot replace external survey validation."

By Behnam Asadi
arXiv Computer Vision
Sep 1

SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation

SNF-Bench is an evaluation framework for long‑horizon fixed‑camera video generation that separates static background fidelity from dynamic flow persistence and drift leakage. It reports these three factors independently, using controlled injections of translation, rotation, scale drift, and progressive freezing to validate each metric’s sensitivity. Auditing public checkpoints shows that whole‑frame motion metrics can mislead, while SNF‑Bench reveals the true trade‑offs between motion quality and background stability.

By Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park
arXiv Computer Vision
Sep 4

SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.

By Shaoliang Yang, Jun Wang
arXiv AI
Jul 7

SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness

arXiv:2607. 02886v1 Announce Type: cross Abstract: Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsely rejecting authentic content, a regime where standard metrics such as AUROC fail to reflect actual operating behavior.

By Jongyeop Hyun, Hyounghun Kim
arXiv Computer Vision
Sep 1

PERSIST: Persistent-State Discrimination for Shot Boundary Detection

PERSIST redefines shot boundary detection as a task of semantic discrimination, requiring a persistent update of a video’s latent temporal state rather than a transient visual change. It employs a FiLM‑conditioned sinusoidal representation network and a structured discriminator that fuses local change, transient impulse, and return‑to‑trend cues into a single interpretable per‑frame signal. The method achieves comparable recall to leading detectors while significantly reducing false positives from flash, text overlay, and archival artifacts, and it is trained solely on real transitions from ClipShots.

By Tingyu Lin, Christian Stippel, Armin Dadras, Jakob Zenzmaier, Florian Kleber, Wolfgang Aigner, Robert Sablatnig
Hugging Face Trending Papers
Aug 19

SPARC: Subspace Position-Aware Robust Few-Shot Calibration for Distribution-Shifted Industrial Anomaly Detection

SPARC is a few‑shot calibration technique for vision‑based industrial anomaly detectors that corrects deployment‑time nuisances by projecting patch features onto a per‑cell subspace, requiring only up to eight verified‑normal images and no gradient updates. It operates between the encoder and detector, using a closed‑form, spatially indexed estimate based on the encoder’s native patch grid. Across seven detectors on shift‑prone benchmarks, SPARC boosts pooled Image AUROC by 13.8 pp and AU‑PRO₀.₃ by 3.5 pp, while showing modest changes on benchmarks without engineered shift.

arXiv Computer Vision
Sep 25

SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range

The SEE Challenge 2026 invites participants to restore RGB images using synchronized event camera data and a target brightness statistic across a wide illumination range. Using the SEE-600K dataset of 610,126 image‑event pairs from 202 real‑world scenes, teams compete under an open‑system protocol, with PSNR as the primary ranking metric and SSIM as a secondary measure. Fifteen valid submissions were evaluated, revealing closely spaced top scores and consistent local errors under severe underexposure, while the report also examines exposure subsets, semantic test cases, shared failure patterns, and system design choices.

By Yunfan Lu, Mingchao Xu, Hanyu Zhou, Shaoyu Liu, Haoyue Liu, Peiqi Duan, Shihan Peng, Yinqiang Zheng, Boxin Shi, Gim Hee Lee, Hui Xiong, Davide Scaramuzza
arXiv AI
Aug 26

A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts

The study audits the AION-1 foundation model, a 39‑modality transformer trained on over 200 million astronomical objects, and finds that its reliance on a survey detection channel—specifically the segmentation map—introduces a severe systematic bias. By keeping image tokens unchanged and editing only the segmentation map, all model outputs (flux, size, ellipticity, redshift) shift by factors of 110–4400 compared to a placebo, revealing that the model’s predictions are driven more by detection gating than by the actual light distribution. This bias propagates into cosmological analyses, shifting tomographic mean redshifts by a median 0.71 × the LSST DESC requirement and exceeding it in multiple assignments, while removing the detection channel eliminates the effect without measurable cost. whyItMatters":"The bias in the detection channel directly inflates errors in key astronomical measurements, potentially compromising the precision of cosmological studies that rely on accurate redshift estimates."

By Ihor Kendiukhov
arXiv Computer Vision
Sep 18

PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos

PROVIA is a system for online mistake detection in egocentric videos that tracks procedure state by maintaining a factual state and a learned summary of performed steps, including mistakes. It uses Bayesian state merging to create an automaton from correct demonstrations and applies a sequential test to convert per‑frame mistake probabilities into alarms. Evaluated on datasets such as CaptainCook4D, IndustReal, HoloAssist, and IMPACT‑ego, PROVIA outperforms baseline methods while operating at 58–70 fps.

By Di Wen, Kailun Yang, Jimmy Weissert, Luc Maria Scherrer, Cedric Z\"ollner, Ruiping Liu, Yufan Chen, Jiale Wei, Junwei Zheng, Kunyu Peng