arXiv Machine Learning

Online Signature Verification Using Augmented Path Signature and T-Mamba

The paper introduces a new online signature verification framework that combines the augmented path signature (APS) descriptor with a T-Mamba model. APS applies time and basepoint augmentations followed by sliding-window path signatures, capturing geometric structures and nonlinear inter-channel interactions. T-Mamba, a hybrid of two temporal convolutional network blocks and a time-scanning Mamba, learns both local temporal patterns and global long-range dependencies, achieving state‑of‑the‑art equal error rates on three public benchmark datasets.

arXiv Machine Learning
Aug 11

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

arXiv:2608. 08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches.

By Pedram MohajerAnsari, Amir Salarpour, Mert D. Pes\'e
arXiv Machine Learning
Sep 10

Rotation-free Online Handwritten Character Recognition Using Linear Recurrent Units

The paper presents a rotation‑free online handwritten character recognition system that uses Sliding Window Path Signature (SW‑PS) to extract local structural features and a lightweight Linear Recurrent Unit (LRU) classifier. The LRU blends the incremental processing of RNNs with the parallel training efficiency of state‑space models to model dynamic stroke characteristics. Experiments on rotated CASIA‑OLHWDB1.1 subsets (digits, English upper letters, Chinese radicals) achieved accuracies of 99.62%, 96.67%, and 94.33% respectively, outperforming competing models in convergence speed and test accuracy.

By Zhe Ling, Sicheng Yu, Danyu Yang
arXiv Machine Learning
Sep 10

AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset

AF-Mamba is a deep learning model that predicts atrial fibrillation (AF) onset one hour in advance using long‑term RR intervals. It combines temporal convolutional networks for local feature extraction with Mamba, a state‑space model for long‑range sequence modeling, achieving high sensitivity (0.889) and specificity (0.943) in subject‑wise testing. The model maintains strong performance across unseen datasets, offering a favorable trade‑off between predictive accuracy and computational efficiency for real‑time ambulatory monitoring.

By Yongbin Lee, Ki H. Chon
arXiv Computer Vision
1d ago

Hybrid Feature Learning for Handwriting Verification

The paper introduces a Hybrid Deep Learning (HDL) architecture that combines Auto-Learned Features (ALF) and Human-Engineered Features (HEF) for handwriting verification. ALF is extracted using a Two Channel Convolutional Neural Network (TC-CNN) or a Two Channel Autoencoder (TC-AE), while HEF is obtained via Gradient Structural Concavity (GSC) or Scale Invariant Feature Transform (SIFT). Experiments on 150,000 pairs of the word "AND" from 1,500 writers show that the HDL model using AE-GSC achieves 99.7% accuracy on a seen writer dataset and 92.16% on a shuffled writer dataset, outperforming CEDAR-FOX, and AE-SIFT performs comparably on unseen writers.

By Seyed Mohammad Abuzar Hashemi, Mihir Chauhan, Jun Chu, Sargur Srihari
arXiv Machine Learning
Aug 11

CPDA: Class-Conditional Path Distribution Alignment for Unsupervised Time-Series Domain Adaptation

arXiv:2608. 09193v1 Announce Type: cross Abstract: Unsupervised time-series domain adaptation (DA) addresses the challenge of transferring a classifier from a labeled source domain to an unlabeled target domain under distribution shifts induced by different users, sensors, devices, acquisition conditions, or temporal dynamics.

By Felix Ott, Christopher Mutschler
arXiv Computer Vision
Aug 26

Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition

The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.

By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin
arXiv Computer Vision
Sep 3

SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

SignMatch introduces a prototype‑structured embedding space that learns to match dictionary sign videos with continuous sign language footage based solely on visual similarity of handshape and motion. By mapping isolated dictionary exemplars into this space, the method enables direct, embedding‑based sign matching and can generalise to unseen signs using only dictionary examples. Experiments on ASL‑Citizen, ChaLearn OSLWL, and BOBSL CSLR2 benchmarks show strong cross‑dataset, cross‑task, and cross‑language performance, outperforming prior approaches on American, British, and Spanish sign languages without benchmark‑specific supervision.

By Ryan Wong, Youngjoon Jang, Liliane Momeni, G\"ul Varol, Andrew Zisserman
arXiv Machine Learning
Aug 14

A Simple State Space Model Excels at Multivariate Time Series Classification

arXiv:2605. 27406v2 Announce Type: replace Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating strong performance through input-dependent state transitions, albeit at considerable complexity.

By Hassan Saadatmand, Geoffrey I. Webb, Hamid Rezatofighi, Mahsa Salehi
arXiv Machine Learning
Jun 24

EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics

arXiv:2606. 24586v1 Announce Type: cross Abstract: Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimization process and the primary evaluation metric, typically the Equal Error Rate (EER).

By Nahuel Gonzalez, Marta Robledo-Moreno, Ivan DeAndres-Tame, Ruben Vera-Rodriguez, Ruben Tolosana