← Back to all news
arXiv AI September 10, 2026 By Ting Liu

Spike-Aware INT8 Execution for Spiking Language Models on Commodity CPUs

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 3

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

arXiv:2606. 03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit.

By Ting Liu
llmsagentsroboticsefficiencybenchmarks
More like this →
arXiv Computation and Language
Sep 10

SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

arXiv:2609.09772v1 Announce Type: new Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spik...

By Ting Liu
efficiency
More like this →
arXiv Machine Learning
Jul 28

FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon

arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.

By Om Mohite
llmsefficiency
More like this →
arXiv Machine Learning
Jul 27

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

arXiv:2607. 22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck.

By Chao Fang, Jun Yin, Man Shi, Marian Verhelst
llmsefficiencybenchmarks
More like this →
arXiv AI
Aug 10

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

arXiv:2604. 26968v2 Announce Type: replace-cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving.

By Sanjeev Rao Ganjihal
agentsefficiency
More like this →
arXiv Machine Learning
Sep 1

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

arXiv:2608.30439v1 Announce Type: cross Abstract: Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space...

By Simon Richter, Ruhai Lin, Jason Yik, Taylor Kergan, Rui-Jie Zhu, Farshad Moradi, Jason Eshraghian
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea