LLM Watermarking as Big Data Provenance: A Deployment-Oriented Systematization
Read the original on arXiv Computation and Language →The paper presents a systematic framework for large language model (LLM) watermarking as a provenance tool in big data ecosystems. It categorizes existing watermarking methods along four deployment dimensions—insertion point, verification authority, operational state, and transformation threat model—and aligns them with the big data principles of Volume, Velocity, Variety, Veracity, and Value. The authors introduce a readiness framework that maps four key workloads—online generation, streaming detection, transformation pipelines, and ecosystem governance—to system-level requirements such as throughput, false-positive control, robustness, cross-domain reliability, governance, and downstream utility, while highlighting gaps between benchmark performance and real-world deployment readiness.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.