arXiv Machine Learning By Alexander Samarin, Sergei Krutikov, Anton Shevtsov, Sergei Skvortsov, Filipp Fisin, Alexander Golubev

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

Read the original on arXiv Machine Learning →

arXiv:2602. 23881v2 Announce Type: replace Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate tokens that are then verified in parallel by the target model.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.