arXiv:2607. 27140v2 Announce Type: replace-cross Abstract: This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic.
By Harvey Samuel George Johnson, Sendy Phang
arXiv:2607. 11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference.
By Wenzong Yang, Danyang Zhang, Kun Cao, Tejus Siddagangaiah, Rajeev Patwari, Zhanxing Pu, Siyin Kong, Zijiang Yang, Hao Zhu, Varun Sharma, Yue Gao, Tianping Li, Fan Yang, Jicheng Chen, Yushan Chen, Fennian Zhao, Aaron Ng, Elliott Delaye, Ashish Sirasao, Sudip Nag
The paper proposes DCO, a dynamic cache orchestration scheme for multi-core AI accelerators that uses application-aware policies and dataflow information to guide cache replacement, bypass decisions, and thrashing mitigation. Using a cycle-accurate simulator, the authors demonstrate up to 1.80× speedup over conventional cache architectures and validate the approach with an analytical model and RTL implementation. The design occupies 0.064 mm² on a 15 nm process and operates at 2 GHz, showing that a shared system-level cache can simplify programming while boosting performance for large language model workloads.
By Zhongchun Zhou, Chengtao Lai, Yuhang Gu, Wei Zhang
Para‑Pipe is a hierarchical mapping framework that combines intra‑ and inter‑stage operator parallelism within a pipelined architecture to optimize deep‑learning inference on heterogeneous System‑on‑Chip (SoC) platforms. By selectively tuning parallelism levels across pipeline stages, it balances throughput and latency while reducing inter‑processor communication overhead. Evaluations on Amlogic and Black Sesame SoCs show Pareto‑optimal configurations, with throughput‑optimized settings achieving up to 11.0 % higher energy efficiency than purely pipelined approaches and 23.3 % over non‑pipelined parallel execution.
Para-Pipe is a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture for machine‑learning computational graphs on heterogeneous System‑on‑Chip (SoC) platforms. By selectively fine‑tuning parallelism levels across pipeline stages, it navigates the trade‑off between throughput and latency, reducing inter‑processor communication overhead and improving energy efficiency. Evaluation on Amlogic and Black Sesame SoCs shows multiple Pareto‑optimal configurations, with throughput‑optimized setups achieving up to 11.0% better energy efficiency than purely pipelined strategies and 23.3% better than non‑pipelined parallel execution.
By Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra
arXiv:2606. 30634v1 Announce Type: new Abstract: Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations leave GPUs idle during pipeline bubbles, wasting computational resources.
By Philip Zmushko, Egor Petrov, Nursultan Abdullaev, Mikhail Khrushchev, Samuel Horv\'ath