arXiv Machine Learning

Watermark Forensics for Generative Models: An Information-Theoretic Perspective

arXiv:2607. 13003v1 Announce Type: cross Abstract: A watermark in a generative model's output is usually asked only whether a text is machine-made.

arXiv Computer Vision
Sep 25

MoSign: Challenge-Response Motion-Watermark Authentication for Anonymous Virtual-Reality Users

MoSign is a challenge-response authentication system that embeds a time‑varying keyed message into the motion of virtual‑reality users, allowing them to prove identity while keeping their avatars anonymous. The watermark is added to the latent space of a motion variational autoencoder and is provably indistinguishable from unwatermarked motion, with security tied to breaking a pseudorandom function. Experiments on HumanML3D and BOXRR‑23 show high authentication accuracy, low false‑accept rates, and resilience against realistic recapture attacks while remaining undetectable by standard detectors.

By Xujun Che, Thomas Carr, Depeng Xu, Aidong Lu, Shuhan Yuan
arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv Computation and Language
Aug 31

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

The paper introduces an adaptive embedding displacement attack (EDA) that exploits rewording, reordering, and resegmentation to remove semantic watermarks from text, achieving a 32.6%–47.9% success rate across four watermarking schemes. To counter this, the authors propose k‑SwordStamp, a semantic watermarking method that uses order‑robust detection over sub‑sentence units, significantly reducing vulnerability to structure‑based edits. Experiments show that EDA remains effective against k‑SwordStamp, but with a lower success rate (10.8%) compared to its performance on other schemes.

By Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum
arXiv AI
Sep 10

PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

The paper introduces PAC‑Private Autoregressive Generation, a method that calibrates noise based on ensemble disagreement across overlapping ‘worlds’ of a private corpus, thereby extending PAC privacy from classification to text generation. By training adapters on a frozen public model and using posterior‑weighted disagreement to add noise only when predictions vary, the approach achieves strong privacy guarantees while preserving most of the fine‑tuning benefit. Experiments on WikiText‑103 with GPT‑2‑small show 74 % of the fine‑tuning gain retained with a per‑token budget of 2⁻³², and membership‑inference success bounded to 51.08 % after one million tokens, outperforming PMixED under matched conditions.

By Mina Mirzadehsarcheshmeh, Amir Keyvan Khandani