arXiv:2605.15508v3 Announce Type: replace
Abstract: The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. This challenge i...
By Jiangnan Yu, Ceyu Xu, Yongji Wu, Yuan Xie
arXiv:2609.36590v1 Announce Type: cross
Abstract: Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeof...
By Hankun Lin, Patrick Pynadath, Ruqi Zhang
arXiv:2608. 15454v1 Announce Type: new Abstract: Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization.
By Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
arXiv:2512.23765v2 Announce Type: replace-cross
Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens and a stronge...
By Tiancheng Su, Meicong Zhang, Guoxiu He
Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains...
arXiv:2609.38510v1 Announce Type: new
Abstract: Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusi...
By Longxuan Yu, Bingsen Chen, Peng Shi, Dongkyu Lee, Yi Xiang, Hideo Kobayashi, Sheng Zhang, Shuaichen Chang, Xing Niu, Zhuoyan Xu, Greg Ver Steeg, Jiarong Jiang