The paper introduces Temporal Residual Neural Radiance Fields for reconstructing dynamic human bodies from monocular video. It builds a temporal residual field independent of MLPs, reduces trainable parameters, speeds up rendering, and employs a multi‑dimensional loss to improve pixel‑level accuracy. Experiments show higher PSNR and SSIM than recent methods while being roughly 780 times faster than Anim‑NeRF and Neural Body.
By Tianle Du, Jie Wang, Xiaolong Xie, Wei Li, Pengxiang Su, Jie Liu
arXiv:2609.14129v1 Announce Type: cross
Abstract: Ultra-High-Definition (UHD) video presents significant challenges for efficient storage and real-time decoding. Learning-based methods, such as Neura...
By Chenhao Zhang, Fengqing Zhu
arXiv:2602.19202v3 Announce Type: replace
Abstract: Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensit...
By Gang Xu, Zhiyu Zhu, Junhui Hou
VOR-Bench is a new benchmark for video object removal that addresses shortcomings in current evaluation methods by providing a dataset with paired edited videos and graffiti masks, a realistic motion-capable paired-video acquisition framework (rMPAF), and a perception-driven scoring model (VOR-MDSM). The dataset includes diverse data from model-generated, tool-rendered, and camera-captured sources, ensuring robust real-world assessment. Experiments show that VOR-Bench’s evaluation results correlate strongly (ρ > 0.9) with human subjective judgments, bridging the gap between traditional metrics and human preference.
By Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du, Zhixiang He, Chi Zhang, Hao Sun, Zhongjiang He, Tianwei Cao, Xuchong Zhang, Hongbin Sun, Kongming Liang, Zhanyu Ma
arXiv:2609.01060v1 Announce Type: cross
Abstract: Compact snapshot hyperspectral cameras provide rich instantaneous spectral measurements for ground-level machine vision, but at lower spatial resolut...
By Mohamad Jouni, Aur\'elien Godet, Mauro Dalla Mura
arXiv:2607. 20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction.
By Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue
arXiv:2609.34895v2 Announce Type: replace
Abstract: Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this li...
By Narges Norouzi, Niccol\`{o} Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus
HyVIC is a configurable variational autoencoder designed for hyperspectral image compression that explicitly separates spatial and spectral feature learning. By allowing independent control of these two aspects, the architecture improves reconstruction fidelity across a wide range of compression ratios, achieving up to 4.66 dB better BD‑PSNR than previous methods. The authors also introduce a metric‑driven strategy for hyperparameter selection and provide code and pretrained models publicly.
By Martin Hermann Paul Fuchs, Behnood Rasti, Beg\"um Demir
arXiv:2507.08375v2 Announce Type: replace
Abstract: Video restoration and enhancement are critical not only for improving visual quality, but also as essential pre-processing steps to boost the perfo...
By Alexandra Malyugina, Yini Li, Joanne Lin, Nantheera Anantrasirichai
TCNeRV is a new implicit neural video compression method that models temporal context in both feature and embedding domains. Its multi‑scale temporal‑context fusion module injects gated historical features across decoder scales, while temporal embedding‑residual coding predicts and encodes only the residual of each content embedding. With about 3 million parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV‑Boost by 2.20 dB and reducing BD‑rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV respectively.
By Xuezhi Xiang, Yixin Zhao, Heqi Xiang, Jiayao Liu, Shanjun Zhang
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.
By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang