AMD + đ¤: Large Language Models Out-of-the-Box Acceleration with AMD GPU
Read the original on Hugging Face Blog âThe Flow has not summarised this story yet â read it at Hugging Face Blog.
The Flow has not summarised this story yet â read it at Hugging Face Blog.
arXiv:2607. 16669v1 Announce Type: cross Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible.
arXiv:2603. 29002v3 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse attention, retrieval-augmented generation (RAG), and compressed contextual memory, to support complex reasoning.
MONA is a new optimizer that extends the Muon optimizer by adding a Nesterovâstyle acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvatureâaware corrections while maintaining Muonâs spectralânorm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on MixtureâofâExperts pretraining across models ranging from 1âŻB to 68âŻB parameters, and achieves stateâofâtheâart performance on downstream benchmarks after fineâtuning the largest model.
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
arXiv:2606. 30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention.