arXiv Machine Learning By Wenjun Qiu, David Lie, Lisa Austin

Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning

Read the original on arXiv Machine Learning →

Calpric is a system that combines automatic text selection, segmentation, active learning, and crowdsourced annotation to create a large, balanced training set for privacy policy classification. By simplifying the labeling task, it enables untrained crowd workers to match the performance of trained annotators and reduces inter‑annotator disagreement, cutting labeling costs. The approach yields a dataset of 16,000 policy text segments across nine data categories and produces models that deliver accurate, fine‑grained labels at a cost of roughly $0.92–$1.71 per segment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.