GaLore: Advancing Large Model Training on Consumer-grade Hardware
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
arXiv:2607. 02558v1 Announce Type: cross Abstract: As machine learning shifts from laboratory curiosity to critical infrastructure, the systems that sustain it span an extraordinary range, from sub-milliwatt microcontrollers to multi-gigawatt datacenter fleets.
arXiv:2606. 24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware.
The paper presents Puro-2B, an open-source language model pretraining recipe that enables training models up to 1.4 trillion tokens on consumer-grade RTX 5090 GPUs using FP8 precision. The authors achieve a best model with a compute cost under $6.9K, approaching Qwen2.5-1.5B performance, and introduce a Puro Cost Scaling Law indicating that about $4.4K suffices to match Qwen2-1.5B. Additionally, they analyze how pretraining data curricula affect downstream performance, providing a full training pipeline and releasing all resources under Apache 2.0.