arXiv Computer Vision

PhysVR: Vision-Language Model Guided Interference-aware Temporal Feature Refinement for Remote Physiological Measurement

arXiv Machine Learning
Aug 4

Rethinking PPG-based Sleep Staging: Datasets, Metrics, and Benchmarks

arXiv:2608. 00943v1 Announce Type: cross Abstract: Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard.

By Shuntian Zheng, Jiawei Wang, Cong Fu, Huan Yu, Chen Chen, Yu Guan, Sai Gu
Hugging Face Trending Papers
Aug 19

EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment

EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.

arXiv Computer Vision
Sep 25

$\unicode{x1F493}$Heartian: Physiology-Aware Relightable Gaussian Head Avatar

The paper introduces Heartian, a physiology‑aware framework that augments Gaussian head avatars with cardiac‑cycle‑dependent albedo modulation, enabling the encoding of remote photoplethysmography (rPPG) signals. By supervising with synchronized contact PPG, the method models the cardiac waveform as a sum of two Gaussians and learns per‑frame spatial residuals via a lightweight MLP. Experiments on 152 stationary recordings from UBFC‑rPPG, PURE, and MMPD show heart‑rate estimation errors as low as 0.29 bpm MAE and 0.38 % MAPE, while preserving reconstruction quality with negligible PSNR loss.

By Xiaoyue Fan, Jose Echevarria, Akshay Paruchuri, Kaan Ak\c{s}it
arXiv AI
Sep 25

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.

By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez