arXiv AI

Dual-Locking Learned AI Models: A PIN-Based Sparse QIM Watermarking and Adaptive Index Permutation Approach

arXiv Machine Learning
Sep 7

A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification

The paper introduces REMARK, a watermark‑based fingerprint framework designed to verify ownership of Graph Neural Networks (GNNs). REMARK generates in‑distribution watermark graphs that maximize output differences between GNN models, thereby reducing performance loss from out‑of‑distribution watermarks. It then extracts robust fingerprints from these output differences, eliminating the need for surrogate models trained on watermark data or reliance on specific output levels, and achieves state‑of‑the‑art verification accuracy across real‑world datasets and GNN architectures.

By Han Zhang, Yan Wang, Guanfeng Liu, Pengfei Ding, Huaxiong Wang, Kwok-Yan Lam
arXiv Machine Learning
Sep 11

CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration

The paper introduces CertDW, a certified dataset watermark and ownership verification method that remains reliable even under malicious perturbations. By leveraging conformal prediction, it defines two statistical measures—principal probability (PP) and watermark robustness (WR)—to evaluate model stability on benign versus watermarked samples. The authors derive certification conditions linking WR to a PP-based threshold and provide a high‑probability bound on false positives, enabling robust ownership verification when a suspicious model’s WR exceeds the PP values of benign models.

By Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao
arXiv Machine Learning
Aug 27

MeMark: Membrane-Space Watermarking for Spiking Neural Networks

MeMark introduces a watermarking scheme for Spiking Neural Networks that embeds a multi‑bit identifier directly into the membrane state of selected Leaky Integrate‑and‑Fire neurons, rather than in the output head. The watermark is recoverable by comparing neuron firing thresholds, eliminating the need for a learned decoder. Experiments on various SNN architectures—including a 215.4M‑parameter SpikeGPT checkpoint—show that all 20 independent 64‑bit keys reliably pass verification under a 51/64 rule, remain robust after fine‑tuning, pruning, quantization, and output‑head replacement, and are not recovered by random keys or adaptive attacks within the tested threat model.

By Roberto Ria\~no, Gorka Abad, Stjepan Picek, Aitor Urbieta
arXiv Computer Vision
6d ago

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

FeatMark is a watermarking framework that protects images from text‑to‑image diffusion model mimicry attacks by embedding small, scene‑consistent micro‑features instead of pixel‑level perturbations. It constructs domain‑specific feature banks, selects executable features, and injects them via mask‑guided concept editing to create highly localized, natural edits. Experiments on VGGFace2, CelebA‑HQ, and WikiArt show FeatMark remains robust against ten strong watermark removal attacks and several adaptive attacks, with minimal impact on perceptual quality and extending to video mimicry scenarios.

By Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu
arXiv AI
Jun 18

Revealing Hidden Vulnerabilities in Autoencoders through Gradient Signal Restoration

arXiv:2505. 03646v5 Announce Type: replace-cross Abstract: Adversarial robustness of deep autoencoders (AEs) has received less attention than that of discriminative models, although their compressed latent representations induce ill-conditioned mappings that can amplify small input perturbations and destabilize reconstructions.

By Chethan Krishnamurthy Ramanaik, Arjun Roy, Tobias Callies, Eirini Ntoutsi
arXiv Machine Learning
1d ago

On the Relationship between Model Quantization and Model Inversion Attacks

The paper investigates how reducing numerical precision through model quantization impacts the vulnerability of neural networks to model inversion attacks. It provides theoretical bounds on mutual information changes and identifies data-dependent effects, especially at 4‑bit precision. Based on these findings, the authors propose a privacy‑aware post‑training quantization strategy that allocates bits adaptively, calibrates activation ranges, and jointly optimizes weight and activation scaling to improve inversion resistance while preserving model utility.

By Rongke Liu, Youwen Zhu
arXiv Machine Learning
Sep 22

On the Information-Theoretic Limits of Latent-Space Watermarking Through Pretrained Generators

The paper investigates latent‑space watermarking using pretrained generators, where a watermark encoder selects latent inputs based on a message and secret key to produce outputs with a specified conditional distribution. For finite alphabets, it derives inner and outer bounds on the rate–key trade‑off and characterizes the capacity region when the generator’s output uniquely determines the latent distribution. The study extends to jointly Gaussian models, identifies key sufficient statistics, optimally allocates secret‑key resources across modes, and analyzes robustness against regeneration attacks, providing compound capacity results and decay rates for repeated attacks.

By Jinwan Jeon, Minju Lee, Sung Hoon Lim
arXiv Computer Vision
Aug 28

Binding Biometrics with AI Agent Identifiers for Delegation of Authority

The paper introduces BIND, a framework that binds a human’s biometric data to an AI agent’s identity and task-specific authority, enabling secure, real‑time delegation of control. By generating a token that an AI agent presents to an Identity Auditor, the system performs biometric authentication and recovers the agent’s ID and scope, providing non‑repudiable proof of human oversight. A practical implementation using face features and a fuzzy commitment scheme with turbo error‑correcting codes achieves a 96% true match rate at zero false match rate and supports 1024‑bit agent tokens.

By Joseph Geo Benjamin, Anil K Jain, Karthik Nandakumar
arXiv AI
3d ago

WARP: A Unified Benchmark for Invisible Image Watermarking -- Robustness and Protection Against Attacks

arXiv:2609.40031v1 Announce Type: cross Abstract: Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated co...

By Khaled Abud, Aleksey Yakushev, Aleksandr Akimenkov, Irina Serzhenko, Kirill Aistov, Egor Kovalev, Dmitry Obydenkov, Sergey Lavrushkin, Anastasia Antsiferova, Dmitriy Vatolin, Yury Markin, Kirill Lukianov