arXiv AI

Infant Care Video Dataset for Classification of Interventions Using Transformers

The Infant Care Video Dataset (ICVD) contains 4,144 videos covering 12 simulated infant care intervention classes, designed to aid automated documentation in neonatal intensive care units. The dataset was collected using a manikin-based setup that varies camera angles and clinician skin tones while maintaining privacy. Baseline experiments with video transformer models (TimeSformer and MotionFormer) achieved over 93% top‑1 accuracy, whereas a framewise approach scored only 23%, highlighting the importance of temporal modeling for this task.

arXiv Computer Vision
Sep 3

Wound3DAssist: A Practical Framework for 3D Wound Assessment

Wound3DAssist is a practical framework that creates 3D wound models from short handheld videos taken with consumer‑grade devices, enabling non‑contact, automatic measurements of wound surfaces. The system integrates 3D reconstruction, wound segmentation, tissue classification, and periwound analysis into a modular workflow. Evaluations on digital models, silicone phantoms, and real patients show millimeter‑scale reconstruction accuracy and multi‑view tissue composition analysis, with full assessments completed in under 20 minutes.

By Remi Chierchia, Rodrigo Santa Cruz, L\'eo Lebrat, Yulia Arzhaeva, Mohammad Ali Armin, Jeremy Oorloff, Chuong Nguyen, Olivier Salvado, Clinton Fookes, David Ahmedt-Aristizabal
arXiv Computer Vision
Sep 3

VideoPulse: Neonatal heart rate and peripheral capillary oxygen saturation (SpO2) estimation from contact free video

VideoPulse is a neonatal dataset and end‑to‑end pipeline that estimates heart rate and peripheral capillary oxygen saturation (SpO2) from facial video without contact. The dataset contains 157 recordings from 52 neonates, and the pipeline uses face alignment, artifact‑aware supervision, and 3D CNN backbones to produce predictions every 2 seconds. On the NBHR dataset the model achieves a heart‑rate MAE of 2.97 bpm and SpO2 MAE of 1.69 %. "whyItMatters":"The results show that short, unaligned neonatal video segments can accurately estimate vital signs, offering a low‑cost, non‑invasive monitoring option for neonatal intensive care."

By Deependra Dewagiri, Kamesh Anuradha, Pabadhi Liyanage, Helitha Kulatunga, Pamuditha Somarathne, Udaya S. K. P. Miriya Thanthrige, Nishani Lucas, Anusha Withana, Joshua P. Kulasingham
arXiv Machine Learning
Jul 20

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

arXiv:2607. 15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods.

By Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
Hugging Face Trending Papers
Jul 6

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.

arXiv Computer Vision
Sep 3

Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation

The paper presents a method for improving markerless 3D pose estimation in infants by cross‑model distillation. Using unannotated infant video, a frozen Sapiens 2 pose model provides dense pseudo‑labels that guide fine‑tuning of the SAM 3D Body model. On a held‑out dataset of eleven infants, the fine‑tuned model shows significant gains in 2D keypoint accuracy and 3D joint error compared to the original SAM 3D Body model.

By R. James Cotton, Divya Joshi, Colleen Peyton
arXiv AI
Aug 25

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

The study presents a vision‑language model pipeline that estimates dynamic, triaxial, bilateral external hand forces during manual material handling tasks using only RGB video and known box mass. By combining text‑guided ROI localization, pretrained vision‑transformer features, and transformer‑based temporal regression, the model achieved root mean square errors of about 4.7–5.6 N for horizontal and mediolateral forces and 10.6–11.0 N for vertical forces across various camera setups. The approach demonstrated that including the handled object as a second ROI and using multi‑camera capture improved peak‑force estimation, showing the feasibility of sensor‑free force estimation for occupational exposure assessment.

By Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum