GaLore: Advancing Large Model Training on Consumer-grade Hardware
Related stories
Intel and Hugging Face Partner to Democratize Machine Learning Hardware Acceleration
AMD + ๐ค: Large Language Models Out-of-the-Box Acceleration with AMD GPU
MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems
arXiv:2607. 02558v1 Announce Type: cross Abstract: As machine learning shifts from laboratory curiosity to critical infrastructure, the systems that sustain it span an extraordinary range, from sub-milliwatt microcontrollers to multi-gigawatt datacenter fleets.
BluTrain: A C++/CUDA Framework for AI Systems
arXiv:2606. 24780v1 Announce Type: new Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware.
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
arXiv:2607. 01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models.
FlowPlace: Flow Matching for Chip Placement
arXiv:2604. 23658v2 Announce Type: replace-cross Abstract: Chip placement plays an important role in physical design.
Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml
arXiv:2602. 15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which could be used in future high-energy physics experiments.
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
arXiv:2605. 10886v3 Announce Type: replace-cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8.
Accelerate Large Model Training using PyTorch Fully Sharded Data Parallel
Predict before you train: Scaling Laws for particle physics foundation models
arXiv:2607. 23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent.
Physics-Distilled Neural Network enabled by Large Language Models for Manufacturing Process-Property Predictive Modeling
arXiv:2606. 11605v1 Announce Type: cross Abstract: Predicting process-property relationships in manufacturing is often challenged by high experimental costs and the limited interpretability of complex 'black-box' models.