arXiv AI

Multi-primitive in-memory computing for Monte Carlo tree search

arXiv:2607. 22869v1 Announce Type: cross Abstract: Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment.

arXiv Machine Learning
Jul 16

Optimizing Binary and Ternary Neural Network Inference on RRAM Crossbars using CIM-Explorer

arXiv:2505. 14303v3 Announce Type: replace-cross Abstract: Using Resistive Random Access Memory (RRAM) crossbars in Computing-in-Memory (CIM) architectures offers a promising solution to overcome the von Neumann bottleneck.

By Rebecca Pelke, Jos\'e Cubero-Cascante, Nils Bosbach, Niklas Degener, Florian Idrizi, Lennart M. Reimann, Jan Moritz Joseph, Rainer Leupers
arXiv Machine Learning
Jul 30

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

arXiv:2607. 26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs.

By Ming-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka, Tushar Krishna, Muhammed Ahosan Ul Karim, Shimeng Yu
arXiv Machine Learning
Sep 4

RACE-AIMC: Selective Inference for Heterogeneous Analog In-Memory Accelerators at the Edge

RACE-AIMC is a framework that selects a single analog in‑memory computing (AIMC) accelerator from a pool to meet a specified energy budget while providing a mathematically exact upper bound on its error rate. Offline, it evaluates each chip, chooses the best one, and computes the bound; online, only that chip runs and a lightweight check decides whether to accept its output or defer to a fallback. Simulations show the certified error stays below 10% (mean 7.83%) and the system achieves clean‑digital accuracy while reducing energy use by about 69% compared to running all chips.

By Osama Yousuf, Martin Lueker-Boden
arXiv AI
6d ago

ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers

ENAS is a hardware‑aware neural architecture search framework tailored for TinyML on microcontrollers. It uses a static feasibility check, a cell‑based search space with various block types and skip connections, and a three‑stage hybrid search strategy (random → top‑K → mutation) with cross‑run caching. The framework runs efficiently without GPUs, achieving significant search‑time speedups and competitive accuracy on Visual Wake Words and Melanoma Cancer benchmarks across a range of microcontrollers.

By Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan
arXiv Machine Learning
Jul 14

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

arXiv:2607. 11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference.

By Wenzong Yang, Danyang Zhang, Kun Cao, Tejus Siddagangaiah, Rajeev Patwari, Zhanxing Pu, Siyin Kong, Zijiang Yang, Hao Zhu, Varun Sharma, Yue Gao, Tianping Li, Fan Yang, Jicheng Chen, Yushan Chen, Fennian Zhao, Aaron Ng, Elliott Delaye, Ashish Sirasao, Sudip Nag
arXiv Machine Learning
Sep 21

Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

The paper reports on programming AMD XDNA NPUs for the FlashAttention workload using open‑source IRON and MLIR‑AIR compiler tools. It compares four reference designs on XDNA 1 and XDNA 2, showing that a fused kernel that keeps QKᵀ scores in local memory achieves 3.62 TFLOP/s on XDNA 2, doubling throughput and greatly improving energy efficiency over the IRON design and the integrated GPU. Roofline analysis guides when to fuse or stream operators based on each device’s ridge points, and the authors release the reference designs as open source.

By Erwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber, Samuel Bayliss