OpenAI Blog

Adversarial training methods for semi-supervised text classification

arXiv Machine Learning
1d ago

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.

By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)
Towards Data Science
Aug 5

Introduction to Semi-Supervised Learning

A primer about Semi-Supervised Learning, the approaches taken with different algorithms and the limitations of using unlabelled data. The post Introduction to Semi-Supervised Learning appeared first on Towards Data Science .

By Carolina Bento
arXiv Machine Learning
Sep 11

Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2

The paper evaluates membership inference attacks (MIAs) on NLP text classifiers using the GLUE SST‑2 sentiment dataset. It compares a TF‑IDF + Logistic Regression pipeline with a fine‑tuned DistilBERT model under a loss‑threshold MIA, finding that both models leak membership signals despite high accuracy. The study also tests mitigations, showing that stronger regularization reduces leakage for Logistic Regression at a utility cost, while fine‑tuning DistilBERT for fewer epochs lowers leakage with minimal accuracy loss.

By William Novak (Minot State University), Muhammad Abusaqer (Minot State University)
Hugging Face Trending Papers
Aug 10

Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training

Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner maximization may further amplify this bias. Existing methods mitigate this issue by correcting class priors or adapting class wise robust supervision, yet they treat each class in isolation and fail to identify which boundaries drive long tailed collapse.

arXiv AI
Aug 11

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

arXiv:2608. 09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale.

By Kevin Thomas, Milosz Kasprzyk, Reuel C Igbokwe Onuigbo, Elliott Pert, Cameron Tovey, Jo\~ao A. Leite, Olesya Razuvayevskaya, Carolina Scarton