We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.
arXiv:2601. 22813v2 Announce Type: replace Abstract: The NVFP4 lower-precision format, supported in hardware by NVIDIA Blackwell GPUs, promises to allow, for the first time, end-to-end fully-quantized pre-training of massive models such as LLMs.
By Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh
SparseDitto is an agentic sparse compilation framework that jointly synthesizes representation, execution schedule, and hardware mapping for sparse matrix computations on GPUs. It uses structural analysis, a learned template-ranking prior, and LLM-guided lowering to generate CUDA code, with target-GPU profiling refining the plan. The framework supports multiple operators such as SpMV, SpMM, and SpGEMM, adapts to different hardware, and achieves significant speedups over cuSPARSE, including up to 146.61× on certain matrices and 3.39× acceleration for full-batch GCN training.
By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
The paper investigates the block-sparse featurizer (BSF), a model that uses small subspaces as atomic units instead of single directions, aiming to capture features on low-dimensional manifolds common in vision. It identifies that BSF still exhibits classic sparse autoencoder failure modes such as feature splitting and composition. The authors propose architectural improvements, notably a Tournament Top‑K selection rule, which markedly reduces feature splitting, and they extend the block concept to a crosscoder framework.
By Alexandru-Iulius Jerpelea, Amith Ananthram
Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity constraint often causes non-negligible accuracy degradation under post-training pruning.
arXiv:2606. 10445v1 Announce Type: new Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup.
By Jaeseong Lee, Seung-won Hwang, Samyam Rajbhandari