π Accelerating LLM Inference with TGI on Intel Gaudi
Related stories
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
arXiv:2606. 17104v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge.
Tile-Level Activation Overlap for Efficient LLM Inference
arXiv:2607. 02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.
Understanding and Coding the KV Cache in LLMs from Scratch
KV caches are one of the most critical techniques for efficient inference in LLMs in production.
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
arXiv:2607. 11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference.
Accelerating Protein Language Model ProtST on Intel Gaudi 2
Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator
Communication-Efficient Verifiable Attention for LLM Inference
arXiv:2606. 16352v1 Announce Type: cross Abstract: Computation integrity of remote large language model (LLM) serving can be questionable.
Introducing AutoRound: Intelβs Advanced Quantization for LLMs and VLMs
Faster assisted generation support for Intel Gaudi
KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
arXiv:2607. 27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task.
Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference
arXiv:2607. 05475v1 Announce Type: cross Abstract: Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency.
