โ† Back to all news
Hugging Face Blog January 30, 2024

Accelerate StarCoder with ๐Ÿค— Optimum Intel on Xeon: Q8/Q4 and Speculative Decoding

Read the original on Hugging Face Blog โ†’

The Flow has not summarised this story yet โ€” read it at Hugging Face Blog.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 2

Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention

arXiv:2602. 07223v2 Announce Type: replace Abstract: Long-context large language model (LLM) inference has become the norm for today's AI applications.

By Yikang Yue, Yuqi Xue, Jian Huang
llmsefficiencybenchmarks
More like this โ†’
Hugging Face Blog
Feb 28, 2024

StarCoder2 and The Stack v2

More like this โ†’
arXiv AI
Jun 9

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control

arXiv:2606. 08382v1 Announce Type: cross Abstract: Low-rank projection has emerged as a promising approach for compressing the KV cache by exploiting hidden-dimension redundancy.

By Priyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang, SeungJae Lee, Jungwook Choi, Mingu Kang
llmsefficiencybenchmarks
More like this โ†’
Hugging Face Blog
Dec 20, 2023

Speculative Decoding for 2x Faster Whisper Inference

multimodal
More like this โ†’
arXiv Machine Learning
Jul 7

Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B

arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.

By Jaeyeon Kim, Jewon Lee, Bo-Kyeong Kim
llmsdiffusionefficiency
More like this โ†’
arXiv Machine Learning
Jul 27

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

arXiv:2607. 22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck.

By Chao Fang, Jun Yin, Man Shi, Marian Verhelst
llmsefficiencybenchmarks
More like this โ†’
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 ยท bb4ee0e