arXiv AI

Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection

arXiv:2608. 14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven.

arXiv Machine Learning
Sep 15

Bridging the Gap in ECG-Based Emotion Recognition: A Unified Evaluation of Deep Learning Models

The paper evaluates deep learning models for electrocardiogram‑based emotion recognition, focusing on generalization across datasets rather than dataset‑specific performance. It introduces two open‑source tools—ARRC for standardized benchmarking and ARDT for inter‑dataset training—to merge three public AER datasets (CUADS, ASCERTAIN, DREAMER) into a more variable benchmark. Using these tools, the authors compare three prominent deep learning architectures and two CNN baselines with hyperparameter tuning and 10‑fold cross‑validation, revealing trade‑offs between accuracy and model complexity and providing a reproducible benchmark for future research.

By Timothy C Sweeney-Fanelli, Ajan Ahmed, Masudul Imtiaz
arXiv AI
Jun 2

Towards a General Intelligence and Interface for Wearable Health Data

arXiv:2605. 22759v2 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging.

By Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
arXiv Computation and Language
Sep 7

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.

By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
arXiv Computation and Language
Aug 28

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.

By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv Machine Learning
Sep 21

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

The study compares temporal deep learning models—Bidirectional LSTM, Temporal Convolutional Network, and Transformer—for physiological emotion recognition using two multimodal wearable datasets, WESAD and EmoWear. Experiments evaluate wrist-only, chest-only, and multimodal sensor configurations with participant-independent leave-one-subject-out cross-validation, and also explore ensembles, sensor ablation, sampling frequency, and saliency analysis. Results show that the best architecture varies by dataset, multimodal sensing consistently outperforms single-site configurations, and a 4 Hz sampling rate offers a cost-effective operating point.

By Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge
arXiv Machine Learning
Sep 14

Explainable Prediction from Mobile Sensing Data through LLM-guided Concept Integration

The paper introduces a Concept-Integrated Transformer (CIT) that uses a pretrained large language model to generate concept abnormality targets with confidence weights, eliminating the need for manual concept annotation. CIT is applied to mobile sensing data from two longitudinal datasets, achieving the highest F1 score on the AFFECT dataset (0.756) and tying for the highest on a PHQ-9 dataset (0.765). The model’s learned concept scores reveal interpretable behavioral and physiological patterns, such as differences in sleep quantity and quality between high and low negative affect groups.

By Yuning Wang, Iman Azimi, Amir M. Rahmani, Pasi Liljeberg
arXiv AI
Jun 15

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

arXiv:2606. 14604v1 Announce Type: cross Abstract: Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking.

By Pavlos Nicolaou, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios