Hugging Face Blog

How to build scalable web apps with OpenAI's Privacy Filter

arXiv Machine Learning
Aug 20

Model Card for OpenAI Privacy Filter

The OpenAI Privacy Filter is a 1.5‑billion‑parameter, bidirectional token‑classification model that detects and redacts personally identifiable information and secrets in unstructured text. It is built from an autoregressive checkpoint, converted into a banded‑attention classifier, and uses a constrained Viterbi decoder to produce coherent spans across eight privacy categories in a single forward pass. The model supports configurable precision‑recall tradeoffs, a 128,000‑token context window, and is designed for efficient local deployment and domain‑specific fine‑tuning as a data‑minimization component within layered privacy workflows.

By Charles de Bourcy, Sahra Ghalebikesabi, Avi Schwarzschild, Alex Gorbachev, Mihai Maruseac, Annie Chu, Vol Kyrylov, Tong Mu, Ally Bennett, Andy Nguyen, Casey Meehan, Jessica Gan Lee, Shane Bauer, Harold Nguyen, Rodolpho Eckhardt, Yuqi Liu, Charlie Oxborough, Marco Rougeth, Omar Chedid, Caio Costa, Yash Parikh, Yao Li, Congzheng Song, Om Thakkar, Vinnie Monaco
OpenAI Blog
Dec 11, 2015

Introducing OpenAI

OpenAI is a non-profit artificial intelligence research company. Our goal is to advance digital intelligence in the way that is most likely to benefit humanity as a whole, unconstrained by a need to generate financial return.