Remote Photoplethysmography (rPPG) enables contactless pulse estimation from facial videos, serving as a vital tool for health monitoring. However, current deep learning methods often struggle under complex disturbances, particularly varying illumination, facial expressions, and unconstrained head movements.
Remote photoplethysmography (rPPG) estimates the blood volume pulse (BVP) signal from facial videos, enabling contact-free health monitoring. Conventional clip-wise approaches, which use video clips as input, require capturing over one hundred frames before inference, thus introducing several seconds of delay and hindering real-time use.
arXiv:2609.12668v1 Announce Type: new
Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
By Chenxi Yang, Yassine Ouzar, Larbi Boubchir
arXiv:2608. 00943v1 Announce Type: cross Abstract: Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard.
By Shuntian Zheng, Jiawei Wang, Cong Fu, Huan Yu, Chen Chen, Yu Guan, Sai Gu
arXiv:2609.05550v1 Announce Type: cross
Abstract: Near-infrared (NIR) video is a promising modality for contactless sleep monitoring, but recent video-based sleep staging methods often use it as a ro...
By Kunmin Jang, You Rim Choi, Hun Heo, Heonjun Lee, Suahn Bae, Dongik Park, Hyun-Woo Shin, Hyung-Sin Kim
arXiv:2606. 12378v1 Announce Type: cross Abstract: Physiological awareness is important for service, social, and assistive robots that interact with humans in everyday environments.
By Zhi Wei Xu, Torbj\"orn E. M. Nordling
EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.
The paper introduces Heartian, a physiology‑aware framework that augments Gaussian head avatars with cardiac‑cycle‑dependent albedo modulation, enabling the encoding of remote photoplethysmography (rPPG) signals. By supervising with synchronized contact PPG, the method models the cardiac waveform as a sum of two Gaussians and learns per‑frame spatial residuals via a lightweight MLP. Experiments on 152 stationary recordings from UBFC‑rPPG, PURE, and MMPD show heart‑rate estimation errors as low as 0.29 bpm MAE and 0.38 % MAPE, while preserving reconstruction quality with negligible PSNR loss.
By Xiaoyue Fan, Jose Echevarria, Akshay Paruchuri, Kaan Ak\c{s}it
arXiv:2606. 07365v1 Announce Type: cross Abstract: Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings.
By Eloy Geenjaar, Vince Calhoun, Scott Daly, Gouthaman KV, Lie Lu, Trisha Mittal, Daniel P. Darcy
Gaussian head avatars typically model intrinsic facial appearance as temporally static, omitting subtle cardiac-induced skin-color variation. We propose $\unicode{x1F493}$Heartian, a physiology-aware...
arXiv:2606. 15284v1 Announce Type: cross Abstract: Photoplethysmography (PPG) plays a central role in wearable health monitoring and clinical decision support.
By Chenyang He, Xinyi Shao, Shun Huang, Bosong Huang, Daoqiang Zhang, Ming Jing, Cheng Ding
UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.
By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez