arXiv Computer Vision

SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range

The SEE Challenge 2026 invites participants to restore RGB images using synchronized event camera data and a target brightness statistic across a wide illumination range. Using the SEE-600K dataset of 610,126 image‑event pairs from 202 real‑world scenes, teams compete under an open‑system protocol, with PSNR as the primary ranking metric and SSIM as a secondary measure. Fifteen valid submissions were evaluated, revealing closely spaced top scores and consistent local errors under severe underexposure, while the report also examines exposure subsets, semantic test cases, shared failure patterns, and system design choices.

Hugging Face Trending Papers
Jun 28

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event cameras provide asynchronous brightness-change observations with high temporal resolution, but prior works often treat voxel channels as an unordered or static feature stack before fusion, rather than explicitly modeling their within-window temporal evolution, weakening the temporal evidence that makes events useful.

Hugging Face Trending Papers
Jul 23

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at room-scale geometry, baked/global illumination, or text-driven generation. We introduce Lumera (Light-aware Unified Engine-native Reconstruction and Assembly), a benchmark and reference pipeline for engine-native, light-aware 3D scene parsing from a single image.

arXiv AI
Sep 10

WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting

WildRelight is the first in-the-wild dataset designed to evaluate single-image relighting models, featuring high-resolution outdoor scenes captured under strictly aligned, temporally varying natural illuminations paired with high-dynamic-range environment maps. The benchmark demonstrates that state-of-the-art models trained on synthetic data suffer severe domain shifts when applied to real-world imagery. Leveraging the dataset’s temporal structure, the authors introduce a physics-guided inference framework combining Diffusion Posterior Sampling with Temporal Sampling-Aware Test-Time Adaptation, enabling synthetic models to self-supervise and align with real-world statistics on-the-fly.

By Lezhong Wang, Mehmet Onurcan Kaya, Siavash Bigdeli, Jeppe Revall Frisvad
arXiv Computer Vision
Sep 7

E-RGB-D: Real-Time Event-Based Perception with Structured Light

The paper introduces E‑RGB‑D, a real‑time event‑based perception system that combines a Digital Light Processing projector with a monochrome event camera to produce RGB‑D data. By projecting structured light and capturing asynchronous brightness changes, the system can detect color and depth for each pixel, achieving a color detection speed of 1400 fps and a depth detection rate of 4 kHz. The approach enables frameless RGB‑D sensing and delivers colorful point clouds without compromising spatial resolution.

By Seyed Ehsan Marjani Bajestani, Giovanni Beltrame
arXiv AI
4d ago

Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement

Consist‑Retinex introduces a one‑step noise‑emphasized consistency training framework for Retinex‑based low‑light image enhancement. It first decomposes images into reflectance and illumination maps using a Retinex Transformer Decomposition Network, then trains two conditional consistency models with a dual objective that blends trajectory consistency and ground‑truth alignment. The method employs adaptive noise‑emphasized fixed‑point sampling to focus supervision near the inference endpoint, achieving state‑of‑the‑art VE‑LOL‑L scores on paired and unpaired low‑light benchmarks while reducing sampling and training costs.

By Jian Xu, Wei Chen, Shigui Li, Delu Zeng, John Paisley, Qibin Zhao
arXiv AI
Jul 13

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

arXiv:2607. 09114v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos.

By Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li
arXiv Computer Vision
Sep 11

Shedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

The paper introduces a benchmark called Shedding Light to evaluate how well generative image models understand and reproduce lighting. The benchmark tests models by asking them to inpaint a simple object, called a light probe, into real photographs and then compares the generated probe to the ground truth to assess lighting direction, colour, and radiance. The authors provide a scalable protocol and open-source code and data for systematic assessment of photometric accuracy in future models.

By Justine Giroux, Jack Oliver Hilliard, Yannick Hold-Geoffroy, Javier Vazquez-Corral, Jean-Fran\c{c}ois Lalonde