AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU
Related stories
OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research
arXiv:2607. 16669v1 Announce Type: cross Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible.
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
arXiv:2603. 29002v3 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse attention, retrieval-augmented generation (RAG), and compressed contextual memory, to support complex reasoning.
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
Little Brains, Big Feats: Exploring Compact Language Models
arXiv:2606. 30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention.
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
Regression Language Models for Code
arXiv:2509. 26476v3 Announce Type: replace-cross Abstract: We study code-to-metric regression: predicting numeric outcomes of code executions, a challenging task due to the open-ended nature of programming languages.
CodegenBench: Can LLMs Write Efficient Code Across Architectures?
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
arXiv:2608. 12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance.
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
arXiv:2607. 08786v1 Announce Type: cross Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge.
The Weight Is Over - Interactive Diffusion on Consumer GPUs
The paper introduces techniques for efficient on-device diffusion-based image generation on consumer GPUs. It presents an embedding translator that reduces weight and latency by mapping a small text encoder into a larger encoder space, a reproducible sweep recipe for balancing speed, quality, and memory, and an interactive editor that achieves sub‑second time‑to‑first‑image on recent GPUs. These contributions aim to broaden the reach of diffusion pipelines to a wide range of client devices.