arXiv Machine Learning By Ruokai Yin, Priyadarshini Panda

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference

Read the original on arXiv Machine Learning →

arXiv:2608. 01536v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly rely on sparsity to reduce inference cost, but most prior work targets a single sparsity source-either weight or activation-and optimizes for batched multi-user inference.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.