arXiv Machine Learning

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

arXiv:2608. 08947v1 Announce Type: cross Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition.

arXiv Computer Vision
Sep 14

Context-Aware Causal Gaze Forecasting for Human-Vehicle Interaction During In-Cabin Tracking Dropouts

The paper introduces the Causal Context-Gated Forecaster (CCGF) for predicting a driver's gaze during dashboard-mounted tracker dropouts. CCGF uses a 60‑frame history of gaze and head pose combined with DINOv3 scene features, and a learned reliability gate adjusts the influence of these inputs as the dropout progresses. Experiments on 2,047 naturalistic driving events show that live scene updates reduce median error by 33% compared to history‑only forecasting, while frozen scene input yields higher error, demonstrating the value of real‑time scene information.

By Shabnam Shabani, Ghazal Farhani
arXiv Computer Vision
Sep 11

Measuring Browser Webcam Gaze Honestly: A Capture-Clock Methodology and Open Reference Implementation

The paper addresses the problem of inaccurate latency reporting in browser-based webcam gaze trackers, which often timestamp samples at emission rather than capture time. It introduces a method that recovers a per-frame capture clock using the browser’s requestVideoFrameCallback API, enabling precise pairing of source frames with inference results or providing a verifiable lower bound when the engine does not expose its pipeline. An open TypeScript implementation and benchmark harness are released, tested on WebGazer and a FaceMesh+KRR pipeline.

By Chi-Sheng Chen, Gabriel A. Brat
Hugging Face Trending Papers
Aug 19

EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment

EgoHRV is a method that estimates heart rate variability (HRV) and heart rate (HR) from the gaze cameras in egocentric headsets. It uses a 3D backbone and a low–high decomposition module to extract the blood volume pulse signal from gaze video, and aligns frequency‑domain representations of contact‑based and camera‑derived signals through cross‑domain pretraining. The approach achieves state‑of‑the‑art accuracy for HR and HRV estimation and, when integrated into EgoExo4D’s proficiency estimator, improves accuracy by 17.8%.