arXiv AI By Nikola Pi\v{z}urica, Matteo Risso, Nikola Milovi\'c, Alessio Burrello, Igor Jovan\v{c}evi\'c, Conor Heins, Miguel de Prado

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

Read the original on arXiv AI →

arXiv:2607. 17855v1 Announce Type: new Abstract: Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 15

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution

The paper introduces a partition-aware scheduling framework for mobile inference on heterogeneous platforms that combines mobile GPUs and multiple CPU core clusters. It jointly optimizes operator partitioning, device assignment, and execution order for static DAGs of operators, such as those in CNNs or vision transformers. An online iterative search approach decomposes large DAGs into stages, targets critical operators, and uses latency predictors to avoid exhaustive profiling, achieving near‑optimal latency with minimal scheduling overhead.

By Zhuojin Li, Marco Paolieri, Leana Golubchik