Using GPT-4 for content moderation
We use GPT-4 for content policy development and content moderation decisions, enabling more consistent labeling, a faster feedback loop for policy refinement, and less involvement from human moderators.
We’re introducing a new model built on GPT-4o that is more accurate at detecting harmful text and images, enabling developers to build more robust moderation systems.
We use GPT-4 for content policy development and content moderation decisions, enabling more consistent labeling, a faster feedback loop for policy refinement, and less involvement from human moderators.
We are introducing a new and improved content moderation tool. The Moderation endpoint improves upon our previous content filter, and is available for free today to OpenAI API developers.
arXiv:2604. 06205v2 Announce Type: replace-cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types.
The paper presents a scalable approach to harmful content moderation on social media by leveraging large language models (LLMs) for few-shot, in-context learning. Experiments across multiple LLMs show that this method outperforms proprietary baselines such as Perspective and OpenAI Moderation, as well as prior few-shot learning techniques, in detecting harmful content. The study also explores the addition of visual cues like video thumbnails to assess multimodal improvements, highlighting the advantages of LLM-based moderation for dynamic and large-scale content filtering.
Discover how SafetyKit leverages OpenAI GPT-5 to enhance content moderation, enforce compliance, and outpace legacy safety systems with greater accuracy .
arXiv:2607. 24898v1 Announce Type: cross Abstract: State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety guidelines to all users.
The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.
The study audits hate‑speech moderation on Twitter (now X) using 540,000 annotated tweets from a full day. Eighty percent of hateful tweets, including violent content, remained online after five months, and removal was only slightly more likely than for non‑hateful tweets, far below the rates for scams or adult content. Automated detection could not reliably classify hate but ranked it highly, allowing human triage; however, current staffing curbed little exposure, while substantial reductions were financially feasible and far below applicable regulatory fines.
arXiv:2502. 16174v4 Announce Type: replace-cross Abstract: Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time.
Developers can now fine-tune GPT-4o with images and text to improve vision capabilities
The paper investigates whether multimodal large language models (MLLMs) can generate and detect realistic multimodal fake news on social media. Using a multi‑agent framework—comprising a story agent, an image agent, and a critic agent—the authors produced over 9,000 paired multimodal news posts across science, health, and entertainment domains. They benchmarked 16 open‑ and closed‑source MLLMs for automated detection and found that most models fall far short of human accuracy, especially in identifying image authenticity, highlighting the need for stronger defenses against social media fake news.
Our latest image generation model is now available in the API via ‘gpt-image-1’—enabling developers and businesses to build professional-grade, customizable visuals directly into their own tools and platforms.