OpenAI Blog

Extracting Concepts from GPT-4

Using new techniques for scaling sparse autoencoders, we automatically identified 16 million patterns in GPT-4's computations.

OpenAI Blog
Mar 14, 2023

GPT-4

We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.

arXiv Machine Learning
Sep 11

SparseDitto: An Agentic Sparse Compilation Framework through Architecture-Aware Synthesis on GPUs

SparseDitto is an agentic sparse compilation framework that jointly synthesizes representation, execution schedule, and hardware mapping for sparse matrix computations on GPUs. It uses structural analysis, a learned template-ranking prior, and LLM-guided lowering to generate CUDA code, with target-GPU profiling refining the plan. The framework supports multiple operators such as SpMV, SpMM, and SpGEMM, adapts to different hardware, and achieves significant speedups over cuSPARSE, including up to 146.61× on certain matrices and 3.39× acceleration for full-batch GCN training.

By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
arXiv Machine Learning
Aug 31

A Deeper Analysis of Block-Sparse Featurizers

The paper investigates the block-sparse featurizer (BSF), a model that uses small subspaces as atomic units instead of single directions, aiming to capture features on low-dimensional manifolds common in vision. It identifies that BSF still exhibits classic sparse autoencoder failure modes such as feature splitting and composition. The authors propose architectural improvements, notably a Tournament Top‑K selection rule, which markedly reduces feature splitting, and they extend the block concept to a crosscoder framework.

By Alexandru-Iulius Jerpelea, Amith Ananthram
arXiv Computation and Language
Aug 31

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

The paper introduces H-Scale, a lightweight post-processing technique for refining per-group scaling factors in NVFP4 quantized large language models. By using a diagonal second-order proxy from calibration activations, H-Scale selects hardware-valid scales that directly target layer output perturbation rather than just weight reconstruction error. Experiments on mainstream LLMs show that H-Scale improves NVFP4 baselines and brings several variants closer to BF16 performance without adding inference overhead.

By Hao Yu, Zheng Li, Dayiheng Liu, Jianwei Zhang
OpenAI Blog
Apr 14, 2025

Introducing GPT-4.1 in the API

Introducing GPT-4. 1 in the API—a new family of models with across-the-board improvements, including major gains in coding, instruction following, and long-context understanding.