The study presents a proof‑of‑concept just‑in‑time adaptive intervention (JITAI) for sleep support that employs an AI agent to analyze 30 days of personal sleep and behavioral data, such as physical activity, smartphone use, and bedtime routines. The agent, running on Home Assistant, reviews the data, evaluates existing reminders, adapts interventions, and records decisions for human review, while limiting reminders to no more than three per day. Initial runs confirmed technical feasibility, successfully completing data review and intervention decisions and saving decision records for future analysis.
By Nick Rezaee, Chelsea Boccagno
BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv:2606. 17767v1 Announce Type: cross Abstract: Personal health data from wearables are typically presented through dashboards of charts and summary statistics, requiring users to actively interpret patterns and implications.
By Nikola Kovacevic, Bastien Husler, Di Zhuang, Rafael Wampfler, Barbara Solenthaler
arXiv:2609.22463v1 Announce Type: new
Abstract: Sleep monitoring using wearable data has shown promise for personal health, yet large language model (LLM)-based summarization and question answering r...
By Yusheng Tan, Running Zhao, Sofia Angel, Ninghui Hao, Ash Arian, Nikita N. Dulin, Jay Lin, Ou Zhu, Faiza Shaik, Xinxing Yang, Bonnie W. Leung, Katie Roster, Arlene Ruiz de Luzuriaga, Kenneth Lee, Alejandra Lastra, Habibul Ahsan, Guihong Wan
arXiv:2608. 06380v1 Announce Type: cross Abstract: Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID).
By Julian Fierrez, Alejandro Pe\~na, Aythami Morales, Ruben Tolosana, Ruben Vera-Rodriguez, Meenakshi Chatterjee, Ahmaniemi Teemu, Wan-Fai Ng, Walter Maetzler, Nikolay V. Manyakov, Jennifer Kudelka, Ralf Reilmann, C. Janneke van der Woude, Kristen Davies, Victoria Macrae, IDEA-FAST Consortium
arXiv:2608.29241v1 Announce Type: new
Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a casca...
By Zachary Ellis, Spencer Hazel, Adam Brandt, Yajie Vera He, Ernest Lim, Jared Joselowitz
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
By Yifan Wang
The study introduces a clinician‑in‑the‑loop benchmark to assess whether large language models can generate evidence‑grounded Brief Hierarchical Taxonomy of Psychopathology (B‑HiTOP) item profiles from multimodal data, including passive sensing, ecological momentary assessment, and questionnaires. Using the GLOBEM dataset, the authors create 14,592 participant‑day instances aligned to 29 B‑HiTOP items across five spectra, and evaluate evidence compatibility rather than diagnostic accuracy. Two‑stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, yielding more conservative score distributions across models, spectra, and evidence settings.
By Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
The paper introduces StratCBT, a new dataset of 9,688 psychological counseling sessions with 256K utterances, each counselor response aligned to one of eight Cognitive Behavioral Therapy (CBT) strategies. It was created by modeling clients’ negative thoughts and generating high‑quality conversations through self‑chat, using realistic sessions for guidance. Experiments show that strategy‑aligned generation improves professional and effective counseling when evaluated with large language model‑simulated clients.
By Zimu Wang, Yiwen Jiang, Xiangyu Zhao, Yaling Shen, Jiahe Liu, Stephanie Fong, Maxmartwell H Cheng, Guilherme C Oliveira, Anh Nguyen, Robert Desimone, Barnaby Nelson, Dominic Dwyer, Zongyuan Ge
arXiv:2605. 17679v2 Announce Type: replace-cross Abstract: Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments when support is relevant, a challenge we term the diary paradox.
By Zhiyuan Wang, Subigya Nepal, Ariful Islam, Indrajeet Ghosh, Xinyu Chen, Katharine E. Daniel, Laura E. Barnes, Philip Chow
arXiv:2605. 22759v2 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging.
By Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
The paper introduces a clinician‑grounded evaluation platform called InterviewPlayground, which uses a memory‑augmented patient simulator to assess AI‑assisted psychiatric intake systems. It supports comparison across different interviewing styles, reduces clinician workload, and measures clinically relevant performance. In a pilot study, a GPT‑based intake interviewer captured more relevant items but made more unfounded inferences and missed safety concerns compared to clinicians.
By King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob Taylor, Ananya Joshi