arXiv:2605. 25796v2 Announce Type: replace-cross Abstract: Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit.
By Jiahao Huo, Wenjie Qu, Yibo Yan, Kening Zheng, Jiaheng Zhang, Xuming Hu, Philip S. Yu, Mingxun Zhou
The paper introduces a dataset watermarking technique that embeds a watermark by increasing the co‑occurrence of randomly selected word pairs through meaning‑preserving local edits. The watermark can be detected solely from generated text with provable false‑positive control, and experiments on four base models and three datasets show reliable detection (p < 0.01) even when the watermarked data constitutes less than 5% of fine‑tuning tokens. Compared to existing methods, the approach better preserves benchmark utility and semantic integrity.
By Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
arXiv:2606. 18430v1 Announce Type: new Abstract: Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited.
By Chih-Duo Hong, Yen-Pang Chen, Fang Yu
arXiv:2602. 09611v2 Announce Type: replace-cross Abstract: Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language models (LVLMs).
By Yue Li, Xin Yi, Dongsheng Shi, Yongyi Cui, Gerard de Melo, Linlin Wang
arXiv:2604. 25860v2 Announce Type: replace-cross Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specific fingerprints.
By Lucio La Cava, Andrea Tagarelli
R-DEIM Net is a 76‑million‑parameter dual‑expert model designed for paraphrase detection that balances accuracy with computational efficiency. It combines an Interaction Expert, which captures token‑level similarity via multi‑scale 2D convolutions and attention, with a Reasoning Expert that generates human‑readable rationales using a Flan‑T5‑small decoder. On the Quora Question Pairs dataset, the model attains 90.07% accuracy and 90.16% F1‑score, matching strong transformer baselines while producing auxiliary rationales.
By Pushp, Vaibhav Prajapati, Himangshu Sarma