arXiv:2607.16196v2 Announce Type: replace
Abstract: Soft plush companions provide a safe and intuitive platform for affective human-robot interaction, but their deformable structure and distributed t...
By Aleksandrs Vali\v{s}evskis, Aleksandrs Okss, Inese T\=i\c{g}ere, Aleksejs Kata\v{s}evs, Dina Bethere, Anete Hofmane, Airisa \v{S}teinberga, Und\=ine Gavri\c{l}enko, Santa Me\c{l}\c{k}e, Lucie Matou\v{s}kova
arXiv:2508. 12435v2 Announce Type: replace-cross Abstract: While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for external sensors.
By Deqing Song, Weimin Yang, Maryam Rezayati, Hans Wernher van de Venn
The study compares temporal deep learning models—Bidirectional LSTM, Temporal Convolutional Network, and Transformer—for physiological emotion recognition using two multimodal wearable datasets, WESAD and EmoWear. Experiments evaluate wrist-only, chest-only, and multimodal sensor configurations with participant-independent leave-one-subject-out cross-validation, and also explore ensembles, sensor ablation, sampling frequency, and saliency analysis. Results show that the best architecture varies by dataset, multimodal sensing consistently outperforms single-site configurations, and a 4 Hz sampling rate offers a cost-effective operating point.
By Desta Haileselassie Hagos, Saurav Keshari Aryal, Legand L. Burge
The paper evaluates deep learning models for electrocardiogram‑based emotion recognition, focusing on generalization across datasets rather than dataset‑specific performance. It introduces two open‑source tools—ARRC for standardized benchmarking and ARDT for inter‑dataset training—to merge three public AER datasets (CUADS, ASCERTAIN, DREAMER) into a more variable benchmark. Using these tools, the authors compare three prominent deep learning architectures and two CNN baselines with hyperparameter tuning and 10‑fold cross‑validation, revealing trade‑offs between accuracy and model complexity and providing a reproducible benchmark for future research.
By Timothy C Sweeney-Fanelli, Ajan Ahmed, Masudul Imtiaz
arXiv:2608. 09830v1 Announce Type: new Abstract: Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders.
By Samaneh Rezaeimanesh, Mohsen Behradfar, Mohammad Fili, Guiping Hu
arXiv:2606. 26723v1 Announce Type: cross Abstract: Respiratory activity is a direct and interpretable physiological channel for wearable stress and affective-state recognition, yet many studies emphasize classification accuracy without identifying which respiratory properties separate different states.
By Andrei Velichko, Mehmet Tahir Huyut
arXiv:2609.14783v1 Announce Type: cross
Abstract: Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, and contact stability, are hard to infer...
By Mashood M. Mohsan, Muhayy Ud Din, Binzhao Xu, Ahmad Abubakar, Irfan Hussain
arXiv:2607. 20820v1 Announce Type: new Abstract: Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive.
By Christian Arzate Cruz, Stefanos Gkikas, Houshyar Asadi
The study investigates respiratory signals from the WESAD dataset to detect stress and other affective states. It compares compact 1‑D CNN models trained on raw 60‑second signals with handcrafted respiratory signatures that capture timing, variability, waveform, spectral, and autocorrelation features. While the CNN achieves the highest accuracy for stress detection, the handcrafted signatures provide stronger, physiologically interpretable markers for baseline, amusement, and especially meditation states.
By Andrei Velichko, Mehmet Tahir Huyut
arXiv:2607. 22779v1 Announce Type: cross Abstract: Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control.
By Federico Del Pup, Elisa Tentori, Manfredo Atzori
This study evaluates machine learning and deep‑learning models for classifying balanced versus imbalanced postural states in immersive virtual reality using a multimodal dataset of kinematic, EMG, and EDA signals. The Mamba‑inspired CNN (MI‑CNN) achieved the highest accuracy (96.76%) and, through SHapley Additive exPlanations (SHAP), identified kinematic features as the most influential for detecting imbalance. Even after reducing input dimensionality by 33% based on SHAP importance, the model maintained near‑optimal performance (0.957 accuracy and F1‑score).
By Nipa Anjum, Md Irfan Pavel, Robert Gonzalez Jr, Kevin Desai, Alberto Cordova, M. Rasel Mahmud, John Quarles
arXiv:2606. 10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals.
By Youcef Soufiane Gheffari, Samiya Silarbi