arXiv Machine Learning

AI-RAN on NPUs: Baseband Processing Without Baseband Chips

arXiv:2607. 04224v1 Announce Type: cross Abstract: AI-RAN aims to unify artificial intelligence and radio access network workloads on a shared compute substrate.

arXiv Machine Learning
Sep 18

Radio-Frequency Convolutional Neural Networks

The paper introduces Radio‑Frequency Convolutional Neural Networks (RF‑CNNs), which repurpose the frequency mixer in wireless radios to perform convolutional neural network inference directly on edge devices. By mapping multi‑channel convolutions onto frequency tones, the passive mixer can execute the entire operation in a single pass, enabling deep CNNs with up to 26.4 million parameters and nine layers to run on smartphones, wearables, and drones. Experimental results show near full‑precision performance while reducing energy consumption to 0.72 fJ per multiply‑accumulate—two orders of magnitude lower than adding a digital processor. "whyItMatters":"The approach leverages existing radio hardware to deliver efficient, state‑of‑the‑art AI inference on billions of devices without increasing size, weight, power, or cost."

By Zhihui Gao, Shi-Yuan Ma, Yiran Chen, Dirk Englund, Tingjun Chen
arXiv Machine Learning
Jun 30

Harvesting AI Computation at the Edge via Generic Approximation

arXiv:2606. 29518v1 Announce Type: cross Abstract: With the widespread adoption of AI in various IoT scenarios such as smart sensing and processing, AI chips have become a common component at the edge.

By Yihan Wang, Huiru Yan, Luxin Zhang, Long Cheng, Weiwei Chen, Ying Wang, Lei Zhang, Cheng Liu, Huawei Li
arXiv Machine Learning
4d ago

Mixture-of-Kittens: MoE Megakernel for NVL72s

arXiv:2609.36070v1 Announce Type: cross Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...

By Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher R\'e
arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim
arXiv AI
Aug 28

Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models

The paper reports the development of 2‑bit microkernels for CPUs and mixed‑precision 2‑bit kernels for Intel Xe2 GPUs, achieving near‑roofline performance. Integrated into LLM inference pipelines, these kernels deliver up to 7× speedup over 16‑bit inference on CPUs and 6.7× on GPUs, surpassing the current state‑of‑the‑art bitnet.cpp runtime by 2.2×. The work demonstrates that ultra‑low‑bit LLM models can be deployed efficiently, offering significant gains in latency, memory, throughput, and energy consumption.

By Evangelos Georganas, Dhiraj Kalamkar, Alexander Heinecke, Pradeep Dubey
arXiv Machine Learning
1d ago

AIR-LLM: Broadcasting AI Weights over Radio for Memory-Free Edge LLM Inference via RF Computing

AIR-LLM is an edge inference architecture that broadcasts large language model (LLM) weights over radio, allowing edge devices to perform matrix-vector multiplications directly in the RF domain without storing or loading the weights. The system uses MIMO spatial multiplexing and an energy‑efficient precoder‑postcoder pair to reduce airtime and calibrate the wireless channel, enabling a single broadcast to serve unlimited users. Experiments on real urban channel models show that AIR-LLM achieves only a 4.0% perplexity loss on LLaMA‑3.1‑8B while saving energy by up to 157.7× compared to FP16 and reducing airtime by over 100× for 20 users.

By Zhihui Gao, Tingjun Chen, Dirk Englund