Releasing Swift Transformers: Run On-Device LLMs in Apple Devices
Related stories
Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms
LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!
Transformers now runs llama.cpp quants
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
arXiv:2607. 00501v1 Announce Type: cross Abstract: We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date.
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
LeanStream introduces a speculate‑and‑refine streaming framework that enables efficient on‑device inference of large language models by progressively refining computation, loading, and cache‑retention priorities using partial GPU results. This approach allows fine‑grained overlap between GPU execution and storage I/O, avoiding the trade‑offs of existing systems that serialize execution or incur redundant weight fetches. Implemented on mobile and embedded platforms, LeanStream reduces memory usage by 4.8× to 7.5× compared to prior work while improving token generation throughput by 1.6× to 2.1×.
A Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026
A Round Up And Comparison of 10 Open-Weight LLM Releases in Spring 2026
Making LLMs lighter with AutoGPTQ and transformers
Graphcore and Hugging Face Launch New Lineup of IPU-Ready Transformers
Faster Stable Diffusion with Core ML on iPhone, iPad, and Mac
Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers
Mitigating Compiler Fusion-Induced Power Bursts in Mobile NPU Inference as the Battery Depletes
arXiv:2607. 16555v1 Announce Type: cross Abstract: Mobile devices increasingly rely on real-time NPU inference for camera and perception workloads.
