AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.
arXiv:2607. 05694v1 Announce Type: cross Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion.
The paper introduces the first watermark designed specifically for diffusion language models (DLMs), which generate tokens in arbitrary order unlike traditional autoregressive models. It overcomes the challenge of missing prior tokens by applying the watermark in expectation over the context and promoting tokens that strengthen the watermark when used as context. Experiments show a >99% true positive rate with minimal quality loss and comparable robustness to existing autoregressive watermarks.
arXiv:2607. 08400v1 Announce Type: cross Abstract: LLM agents reach users through resellers, who may rebrand a developer's agent or substitute a cheaper model.
arXiv:2607. 20435v1 Announce Type: cross Abstract: Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights.
FeatMark is a watermarking framework that protects images from text‑to‑image diffusion model mimicry attacks by embedding small, scene‑consistent micro‑features instead of pixel‑level perturbations. It constructs domain‑specific feature banks, selects executable features, and injects them via mask‑guided concept editing to create highly localized, natural edits. Experiments on VGGFace2, CelebA‑HQ, and WikiArt show FeatMark remains robust against ten strong watermark removal attacks and several adaptive attacks, with minimal impact on perceptual quality and extending to video mimicry scenarios.