Hugging Face Trending Papers

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices. In this paper, we present a hardware-oriented methodology for accelerating discrete Bayesian inference on commercial off-the-shelf embedded GPUs.

arXiv AI
Jul 21

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

arXiv:2607. 17855v1 Announce Type: new Abstract: Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices.

By Nikola Pi\v{z}urica, Matteo Risso, Nikola Milovi\'c, Alessio Burrello, Igor Jovan\v{c}evi\'c, Conor Heins, Miguel de Prado
arXiv Machine Learning
Sep 15

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution

The paper introduces a partition-aware scheduling framework for mobile inference on heterogeneous platforms that combines mobile GPUs and multiple CPU core clusters. It jointly optimizes operator partitioning, device assignment, and execution order for static DAGs of operators, such as those in CNNs or vision transformers. An online iterative search approach decomposes large DAGs into stages, targets critical operators, and uses latency predictors to avoid exhaustive profiling, achieving near‑optimal latency with minimal scheduling overhead.

By Zhuojin Li, Marco Paolieri, Leana Golubchik
arXiv Machine Learning
Sep 16

High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models

The paper introduces a tensor-based formulation of the Viterbi algorithm for Hidden Semi-Markov Models (HSMMs), converting inner loops into tensor operations that align with SIMD and massively parallel architectures. It presents optimized implementations for single- and multi-core CPUs and, for the first time, GPUs. Experiments show speedups of up to 14× on a single core, over 200× with multi-core, and more than 570× on GPU compared to the sequential baseline, setting a new performance benchmark for large-scale HSMM decoding.

By Lorenzo Piarulli, Elia Belli, Daniele De Sensi