arXiv Machine Learning

High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models

The paper introduces a tensor-based formulation of the Viterbi algorithm for Hidden Semi-Markov Models (HSMMs), converting inner loops into tensor operations that align with SIMD and massively parallel architectures. It presents optimized implementations for single- and multi-core CPUs and, for the first time, GPUs. Experiments show speedups of up to 14× on a single core, over 200× with multi-core, and more than 570× on GPU compared to the sequential baseline, setting a new performance benchmark for large-scale HSMM decoding.

arXiv AI
Jul 21

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

arXiv:2607. 17855v1 Announce Type: new Abstract: Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices.

By Nikola Pi\v{z}urica, Matteo Risso, Nikola Milovi\'c, Alessio Burrello, Igor Jovan\v{c}evi\'c, Conor Heins, Miguel de Prado
arXiv Machine Learning
Jul 9

VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination

arXiv:2604. 09558v2 Announce Type: replace-cross Abstract: With the widening gap between compute and memory operation latencies, data movement optimizations have become increasingly important for DNN compilation.

By Muyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta, Taeksang Kim, Xin Xu, Janardhan Kulkarni, Ofer Dekel, Vikram Adve, Charith Mendis
arXiv AI
Aug 18

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.

By Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
arXiv Machine Learning
Sep 14

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

The paper investigates how a single SM utilization metric can misrepresent the true workload of large language model (LLM) inference on Nvidia Hopper GPUs. By profiling vLLM with FlashAttention‑3 and cuBLASLt on an H100 NVL across various phases (cold prefill, warm prefill, and decode) and varying sequence length and batch size, the authors replace the single utilization figure with eight detailed counter‑validated views. These views, tied to specific Nsight Compute counters or formulas, reveal how factors such as fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection create utilization gaps across four production models and six per‑layer kernel roles.

By Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic, Marco Chiesa
arXiv AI
Aug 10

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

arXiv:2608. 06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design.

By Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram