Introducing the AMD 5th Gen EPYC™ CPU
Related stories
Hugging Face and AMD partner on accelerating state-of-the-art models for CPU and GPU platforms
AMD and OpenAI announce strategic partnership to deploy 6 gigawatts of AMD GPUs
AMD and OpenAI have announced a multi-year partnership to deploy 6 gigawatts of AMD Instinct GPUs, beginning with 1 gigawatt in 2026, to power OpenAI’s next-generation AI infrastructure and accelerate global AI innovation.
AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU
Creating custom kernels for the AMD MI300
Case Study: Millisecond Latency using Hugging Face Infinity and modern CPUs
Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
The article discusses several recent developments in AI and computing, including the lack of legal rights for machines, the use of SPADE to automate environment generation, and the creation of improved GPU kernels with Hawkeye. It also touches on the differential acceleration of cyber, math, and AI technologies. These topics illustrate ongoing progress in AI infrastructure and performance optimization.
Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study
The paper reports on programming AMD XDNA NPUs for the FlashAttention workload using open‑source IRON and MLIR‑AIR compiler tools. It compares four reference designs on XDNA 1 and XDNA 2, showing that a fused kernel that keeps QKᵀ scores in local memory achieves 3.62 TFLOP/s on XDNA 2, doubling throughput and greatly improving energy efficiency over the IRON design and the integrated GPU. Roofline analysis guides when to fuse or stream operators based on each device’s ridge points, and the authors release the reference designs as open source.
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
arXiv:2606. 11357v1 Announce Type: cross Abstract: With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets.
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
arXiv:2607. 19438v1 Announce Type: cross Abstract: Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API.
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
AsmEvo is an agentic assembly-level optimizer that targets compiled AMDGPU code objects, reconstructing a reassemblable representation and applying low-level edits guided by a long-horizon agent. It rebuilds ABI-preserving optimized objects and verifies functional equivalence through differential testing against the original binary. Experiments show significant speedups—up to 1.35× geometric mean on MI308X and 1.18× on MI300X—while maintaining correctness.
