Steal the Knowledge, Inherit the Mark: TwinMark for Distillation Watermarking via Second Moments and Class Multiplexing
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
TwinMark is a watermarking scheme that embeds a single SHAKE128 secret into a vision model using two complementary linear functionals of model-output summaries: a covariance projector (cov‑Feat) and a class‑conditional Fisher‑aligned linear carrier (cc‑FALC). These readouts cover both classifier APIs attacked by KL knowledge distillation and representation‑only hosts attacked by feature‑matching distillation, each providing a teacher‑measurable a posteriori certificate. Across 13 attacks on datasets such as CIFAR‑10, CIFAR‑100, and Mini‑ImageNet, TwinMark remains detectable on every post‑attack model that retains task utility, survives cross‑architecture distillation onto ResNet‑18/50, VGG‑16, and MobileNet‑V3, and can be ported to GNSS few‑shot, VOC detection, ISIC segmentation, and STL‑10 SimCLR.
arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
arXiv:2608.29030v1 Announce Type: new Abstract: In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus e...
The study investigates what knowledge a student model inherits from its teachers beyond accuracy when using knowledge distillation for encrypted‑traffic classification. By distilling a 101k‑parameter student from two teachers of equal accuracy but different construction, the authors test ten hypotheses over a year of real TLS traffic, finding that unknown‑traffic detection and shortcut reliance can transfer depending on temperature settings and model size, while other abilities do not. The results show that distillation can propagate teacher habits, but some inherited capabilities can also be achieved without a teacher.
MeMark introduces a watermarking scheme for Spiking Neural Networks that embeds a multi‑bit identifier directly into the membrane state of selected Leaky Integrate‑and‑Fire neurons, rather than in the output head. The watermark is recoverable by comparing neuron firing thresholds, eliminating the need for a learned decoder. Experiments on various SNN architectures—including a 215.4M‑parameter SpikeGPT checkpoint—show that all 20 independent 64‑bit keys reliably pass verification under a 51/64 rule, remain robust after fine‑tuning, pruning, quantization, and output‑head replacement, and are not recovered by random keys or adaptive attacks within the tested threat model.
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.