arXiv Machine Learning By Shahir M A

How Weight Encoding Affects Language Model Placement and Performance on the Apple Neural Engine

Read the original on arXiv Machine Learning →

The study examines how different weight encodings—dense fp16, int8, and ternary with two‑bit lookup tables—affect the placement and performance of language models on Apple’s Neural Engine (ANE) via Core ML. Using five checkpoints across two architectures, the authors combine compiler plans, memory‑controller measurements, and compute‑unit controls to assess a single‑token forward workload. Results show that fp16 models may run on the CPU or ANE depending on size, while compressed int8 models consistently activate the ANE and halve warm‑forward latency, demonstrating that encoding influences both placement and speed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Optimizing AI Inference Across the Deployment Stack

The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.

By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel
Hugging Face Trending Papers
Jul 27

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU.

arXiv AI
Jul 13

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

arXiv:2607. 09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs).

By Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini
arXiv Machine Learning
Sep 25

Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhone

The paper introduces Routide, a Swift/MLX runtime that runs a quantized Qwen3.6-35B-A3B model on iPhone by keeping expert weights on device storage and a byte‑budgeted subset in memory. It evaluates cache‑policy effects, showing that a 512 MiB LRU cache yields 0.00% demand hits while a 576 MiB LRU reaches 38.58% hits across five 128‑token workloads, indicating that capacity limits depend on policy and workload. The study also reports memory footprints, thermal events, and power estimates, demonstrating that flash‑backed MoE inference is feasible within bounded resources but has measurable limitations.

By Musa Shams
arXiv AI
Jul 24

Profiling Lightweight Large Language Models

arXiv:2607. 20806v1 Announce Type: new Abstract: Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resource-constrained edge and mobile environments.

By Tomohiro Harada, Enrique Alba, Gabriel Luque