GaLore: Advancing Large Model Training on Consumer-grade Hardware
Related stories
Intel and Hugging Face Partner to Democratize Machine Learning Hardware Acceleration
AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU
MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems
arXiv:2607. 02558v1 Announce Type: cross Abstract: As machine learning shifts from laboratory curiosity to critical infrastructure, the systems that sustain it span an extraordinary range, from sub-milliwatt microcontrollers to multi-gigawatt datacenter fleets.
BluTrain: A C++/CUDA Framework for AI Systems
arXiv:2606. 24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware.
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
The paper presents Puro-2B, an open-source language model pretraining recipe that enables training models up to 1.4 trillion tokens on consumer-grade RTX 5090 GPUs using FP8 precision. The authors achieve a best model with a compute cost under $6.9K, approaching Qwen2.5-1.5B performance, and introduce a Puro Cost Scaling Law indicating that about $4.4K suffices to match Qwen2-1.5B. Additionally, they analyze how pretraining data curricula affect downstream performance, providing a full training pipeline and releasing all resources under Apache 2.0.
Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts
The paper introduces a hardware-software co‑design framework that compresses Mixture‑of‑Experts (MoE) model weights into low‑precision, hardware‑native sparse representations, enabling efficient execution on Sparse Tensor Cores (SpTCs). By relaxing discrete support selection through continuous reparameterization, the method jointly optimizes quantized weights and supports a router‑weighted reconstruction objective, achieving up to 4.35 percentage‑point gains in joint sparse‑quantization accuracy while retaining 96.09% of the original model’s performance. A custom grouped sparse GEMM kernel further boosts inference speed, outperforming NVIDIA’s baseline by up to 1.65× and reducing latency by up to 4.03× on B200 GPUs.
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.
FlowPlace: Flow Matching for Chip Placement
arXiv:2604. 23658v2 Announce Type: replace-cross Abstract: Chip placement plays an important role in physical design.
Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml
arXiv:2602. 15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physics experiments.
A Rapid Pipeline for Training and Deploying ML Models on WeBe Band
The paper presents a rapid pipeline for training and deploying machine‑learning models on the WeBe Band, a wrist‑worn wearable device. It automates the creation of hardware‑efficient models, integrates with the Piccolo AI ecosystem, and supports OTA deployment while profiling latency and memory usage. Experimental results show trade‑offs between classical models and lightweight neural networks for real‑time performance on a microcontroller.
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
arXiv:2605. 10886v3 Announce Type: replace-cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8.