Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed for natural images may disrupt the fine-grained topology and textures essential for identity discrimination.
The paper presents lightweight architectures for detecting GAN-generated synthetic faces, comparing a compact Swin Transformer, pre‑trained Swin‑Tiny and Swin‑Small models, and a hybrid EfficientNet‑B0 + Swin Transformer. Using the 140K Real and Fake Faces dataset, the hybrid model achieved 99% accuracy and 99.44% recall on 5,000 test images, outperforming both pure Swin variants and a CNN‑only baseline. The study demonstrates that combining hierarchical CNN features with shifted‑window self‑attention yields an efficient, computationally lightweight detection method.
By Sejuti Basu, Ashima Sood, Vijay Kumar, Sahil Sharma
arXiv:2609.39585v1 Announce Type: new
Abstract: AI video generators have not only become harder to detect but are used to generate a diverse set of scenarios from landscapes to street views to animal...
By Sidharth Shanu, Gautam Kumar, Tej Singh
arXiv:2606. 08033v1 Announce Type: cross Abstract: Cracks are a critical indicator of building health, and early stage identification is fundamental to prevent harmful damages.
By Mattia Forlesi, Alfonso Esposito, Ivan Zyrianoff, Alessandro Marzani, Marco Di Felice
arXiv:2607. 15288v1 Announce Type: cross Abstract: Facial expression recognition is an important computer vision task with applications in human--computer interaction, mental health monitoring, driver alert systems, and behavioral analysis.
By Chethiya Galkaduwa
arXiv:2607. 22808v1 Announce Type: cross Abstract: The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies.
By Md. Ajwad Hossain
arXiv:2609.14654v1 Announce Type: new
Abstract: Convolutional neural networks (CNNs) are typically evaluated using held-out classification accuracy, an approach that presupposes predictions are based...
By Abhilekha Dalal, Michael Okonoda, Eder Martinez, Lior Shamir
arXiv:2605. 09697v3 Announce Type: replace-cross Abstract: In many real-world computer vision applications, including medical imaging and industrial inspection, binary classification tasks are characterized by a severe scarcity of positive samples.
By Radhika Amar Desai, Modigari Narendra
arXiv:2607. 16196v1 Announce Type: new Abstract: Soft, sensorized companions offer a physically safe and emotionally intuitive interface for socially assistive technologies, yet their deformability and multichannel tactile sensing complicate the robust interpretation of human affect.
By Aleksandrs Vali\v{s}evskis, Aleksandrs Okss, Inese T\=i\c{g}ere, Aleksejs Kata\v{s}evs, Dina Bethere, Anete Hofmane, Airisa \v{S}teinberga, Und\=ine Gavri\c{l}enko, Santa Me\c{l}\c{k}e, Lucie Matou\v{s}kov\'a
arXiv:2307. 00919v2 Announce Type: replace-cross Abstract: Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification benchmarks.
By Vinoth Nandakumar, Arush Tagade, Tongliang Liu
As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge inside them remains largely unexplored.
The paper introduces MoE-JEPA, a dual‑stream deepfake detection model that combines a V‑JEPA backbone with a Residual Mixture‑of‑Experts mechanism and a noise stream branch. It further incorporates a Gated Attention Multiple Instance Learning module to refine spatial semantic understanding. On the SID‑Set benchmark, MoE‑JEPA achieves a new state‑of‑the‑art accuracy of 95.54%, outperforming much larger models.
By Simone Teglia, Irene Amerini