Hugging Face Blog

GaLore: Advancing Large Model Training on Consumer-grade Hardware

arXiv AI
Jun 24

BluTrain: A C++/CUDA Framework for AI Systems

arXiv:2606. 24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware.

By Adhitya Charan, Adwaid Suresh, Anuj Kumar, Aparna A, Dhanakumar K, Dharun M S, Dinesh G, Goutham Kumar Reddy K, Harshini V M, Jenifa D, Jona Delcy C A, Kathirvel S, Killi Uma Maheswara Rao, Kiruthik Kanna M, Kurra Vishnu Sai, Madhumithaa G K, Navin Kumar V, Ram Charan Golla, Revathi T, Rishikkanth R, Sanjay Krishna M V, Surendra Vendra
arXiv Machine Learning
Aug 28

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

The paper presents Puro-2B, an open-source language model pretraining recipe that enables training models up to 1.4 trillion tokens on consumer-grade RTX 5090 GPUs using FP8 precision. The authors achieve a best model with a compute cost under $6.9K, approaching Qwen2.5-1.5B performance, and introduce a Puro Cost Scaling Law indicating that about $4.4K suffices to match Qwen2-1.5B. Additionally, they analyze how pretraining data curricula affect downstream performance, providing a full training pipeline and releasing all resources under Apache 2.0.

By Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao, Yanmohan Wang, Mingzhe Zhang, Kaiyue Wen, Kaifeng Lyu, Wenguang Chen
arXiv AI
1d ago

Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts

The paper introduces a hardware-software co‑design framework that compresses Mixture‑of‑Experts (MoE) model weights into low‑precision, hardware‑native sparse representations, enabling efficient execution on Sparse Tensor Cores (SpTCs). By relaxing discrete support selection through continuous reparameterization, the method jointly optimizes quantized weights and supports a router‑weighted reconstruction objective, achieving up to 4.35 percentage‑point gains in joint sparse‑quantization accuracy while retaining 96.09% of the original model’s performance. A custom grouped sparse GEMM kernel further boosts inference speed, outperforming NVIDIA’s baseline by up to 1.65× and reducing latency by up to 4.03× on B200 GPUs.

By Kwanhee Lee, Namhoon Lee, Dan Alistarh
arXiv Machine Learning
Aug 3

Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

arXiv:2602. 15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physics experiments.

By Katya Govorkova, Julian Garcia Pardinas, Vladimir Loncar, Victoria Nguyen, Sebastian Schmitt, Marco Pizzichemi, Loris Martinazzoli, Eluned Anne Smith
arXiv AI
Sep 25

A Rapid Pipeline for Training and Deploying ML Models on WeBe Band

The paper presents a rapid pipeline for training and deploying machine‑learning models on the WeBe Band, a wrist‑worn wearable device. It automates the creation of hardware‑efficient models, integrates with the Piccolo AI ecosystem, and supports OTA deployment while profiling latency and memory usage. Experimental results show trade‑offs between classical models and lightweight neural networks for real‑time performance on a microcontroller.

By Ehsan Kourkchi, Asmita Asmita, Houman Homayoun, Mahdi Eslamimehr
arXiv AI
Jul 10

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

arXiv:2605. 10886v3 Announce Type: replace-cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8.

By Liang Luo, Yinbin Ma, Quanyu Zhu, Vasiliy Kuznetsov, Yuxin Chen, Neng Shi, Jian Jiao, Jiecao Yu, Buyun Zhang, Tongyi Tang, Xiaohan Wei, Yanli Zhao, Zeliang Chen, Yuchen Hao, Venkatesh Ranganathan, Sandeep Parab, Yantao Yao, Maxim Naumov, Chunzhi Yang, Shen Li, Ellie Wen, Wenlin Chen, Santanu Kolay, Chunqiang Tang