arXiv:2609.00378v1 Announce Type: new
Abstract: Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadrat...
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2608.27782v1 Announce Type: cross
Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is t...
By Xujun Che, Depeng Xu, Shuhan Yuan
arXiv:2607. 08400v1 Announce Type: cross Abstract: LLM agents reach users through resellers, who may rebrand a developer's agent or substitute a cheaper model.
By Zheng Gao, Xiaoyu Li, Xiaoyan Feng, Jiaojiao Jiang, Yang Song, Yulei Sui, Zhenchang Xing, Liming Zhu
MoSign is a challenge-response authentication system that embeds a time‑varying keyed message into the motion of virtual‑reality users, allowing them to prove identity while keeping their avatars anonymous. The watermark is added to the latent space of a motion variational autoencoder and is provably indistinguishable from unwatermarked motion, with security tied to breaking a pseudorandom function. Experiments on HumanML3D and BOXRR‑23 show high authentication accuracy, low false‑accept rates, and resilience against realistic recapture attacks while remaining undetectable by standard detectors.
By Xujun Che, Thomas Carr, Depeng Xu, Aidong Lu, Shuhan Yuan
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.
By Chih-Duo Hong, Yen-Pang Chen, Fang Yu
arXiv:2607. 00224v1 Announce Type: cross Abstract: Watermarking promises a statistical trace of large language model (LLM) use, but real documents, after editing or paraphrasing, rarely arrive as purely human-written or purely machine-generated.
By Shuwen Chai, Qiaosen Wang