OpenAI Blog

Understanding the source of what we see and hear online

Today we’re introducing new technology to help researchers identify content created by our tools and joining the Coalition for Content Provenance and Authenticity Steering Committee to promote industry standards.

arXiv AI
Sep 4

Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

The paper introduces Provenance Density, an interface that visualizes the density of verified claims within a text to counter the Fluency Trap—where users mistake fluent AI-generated hallucinations for truth. In a study with 81 participants, the interface significantly improved users’ ability to distinguish true from fabricated content, while no signal led to no discernment. A technical audit of 200 samples revealed that retrieval density alone is insufficient, and that the Consistency Veto provides most of the discriminative power for dynamic queries.

By Qing Zhang, Yifei Huang, Juyoung Lee, Thad Starner, Jun Rekimoto
arXiv Computation and Language
6d ago

Epstein Files Engine: Agentic Search for Investigative Journalism

The Epstein Files Engine is an AI agent developed by the New York Times to help journalists investigate a massive mixed‑media collection released by the U.S. Department of Justice on January 30, 2026, which contains about three million pages of PDFs related to Jeffrey Epstein. The Engine translates reporter questions into Google BigQuery SQL queries across three corpora—Epstein‑related releases, the Times’s archive, and external Epstein‑related news headlines—using an LLM to plan queries and return citation‑rich answers that reporters can verify. Over 100 journalists used the Engine, contributing to at least 20 published stories, and the system includes a Diff method for text‑and‑visual duplicate matching to surface genuinely new information.

By Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward