arXiv:2607. 17855v1 Announce Type: new Abstract: Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices.
By Nikola Pi\v{z}urica, Matteo Risso, Nikola Milovi\'c, Alessio Burrello, Igor Jovan\v{c}evi\'c, Conor Heins, Miguel de Prado
arXiv:2603. 29002v3 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse attention, retrieval-augmented generation (RAG), and compressed contextual memory, to support complex reasoning.
By Zifan He, Rui Ma, Yizhou Sun, Jason Cong
The paper introduces a partition-aware scheduling framework for mobile inference on heterogeneous platforms that combines mobile GPUs and multiple CPU core clusters. It jointly optimizes operator partitioning, device assignment, and execution order for static DAGs of operators, such as those in CNNs or vision transformers. An online iterative search approach decomposes large DAGs into stages, targets critical operators, and uses latency predictors to avoid exhaustive profiling, achieving near‑optimal latency with minimal scheduling overhead.
By Zhuojin Li, Marco Paolieri, Leana Golubchik
arXiv:2607. 08786v1 Announce Type: cross Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge.
By Tao Lu, Haoyu Wang, Zonghui Wang, Keshen Xiang, Jiaheng Zhang, Wenzhi Chen
arXiv:2501.10375v3 Announce Type: replace-cross
Abstract: Mixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory...
By Yujie Zhang, Shivam Aggarwal, Tulika Mitra
arXiv:2607. 10183v1 Announce Type: cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory.
By Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li
arXiv:2608.28652v1 Announce Type: new
Abstract: Artificial intelligence (AI) models have demonstrated remarkable capabilities across various domains, yet their widespread deployment is impeded by sig...
By Venkat R. Dasari, Jakob A. Adams, Vinod K. Mishra, Brian Jalaian
arXiv:2607. 21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost.
By Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
The paper introduces a tensor-based formulation of the Viterbi algorithm for Hidden Semi-Markov Models (HSMMs), converting inner loops into tensor operations that align with SIMD and massively parallel architectures. It presents optimized implementations for single- and multi-core CPUs and, for the first time, GPUs. Experiments show speedups of up to 14× on a single core, over 200× with multi-core, and more than 570× on GPU compared to the sequential baseline, setting a new performance benchmark for large-scale HSMM decoding.
By Lorenzo Piarulli, Elia Belli, Daniele De Sensi
arXiv:2609.38090v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-...
By Sanjali Yadav, Bahar Asgari
arXiv:2608.30439v1 Announce Type: cross
Abstract: Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space...
By Simon Richter, Ruhai Lin, Jason Yik, Taylor Kergan, Rui-Jie Zhu, Farshad Moradi, Jason Eshraghian
arXiv:2608. 05033v1 Announce Type: cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning.
By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding