arXiv AI

WaveVerif: Acoustic Side-Channel based Verification of Robotic Workflows

arXiv:2510. 25960v2 Announce Type: replace-cross Abstract: In this paper, we present a framework that uses acoustic side-channel analysis (ASCA) to monitor and verify whether a robot correctly executes its intended commands.

arXiv AI
Aug 18

Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification

arXiv:2608. 15636v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment.

By Chunyu Qi, Zhuoran Song, Jian Weng, Haozhe Jiang, Xueyuan Liu, Naifeng Jing, Guanghui He, Xiaoyao Liang, Haibing Guan
arXiv Machine Learning
Sep 24

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.

By Julia Huang
arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv AI
Oct 1

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study. "whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."

By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
arXiv AI
Jun 30

Tactile Gesture Recognition with Built-in Joint Sensors for Industrial Robots

arXiv:2508. 12435v2 Announce Type: replace-cross Abstract: While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for external sensors.

By Deqing Song, Weimin Yang, Maryam Rezayati, Hans Wernher van de Venn
arXiv Machine Learning
Sep 22

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.

By Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
arXiv AI
Jun 9

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.

By Aueaphum Aueawatthanaphisut
arXiv Machine Learning
Jun 8

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

arXiv:2606. 06837v1 Announce Type: cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself.

By Vsevolod (V.), Kovalev, Pranay Manocha