arXiv Computation and Language By Jaeduk Lee, Wan Choi

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

Read the original on arXiv Computation and Language →

The paper introduces X-CoSD, a communication‑efficient framework for collaborative speculative decoding that allows a small on‑device language model to draft tokens while a server‑side large language model verifies them, even when the two models use different vocabularies. X-CoSD employs hybrid resampling to limit distribution exchange to only the common vocabulary, and its enhanced variant X-CoSD‑E further reduces communication by sending only server‑sampled replacement candidates for device verification. Both variants preserve the server model’s distribution and, according to experiments, accelerate token generation without sacrificing quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 7

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.

By Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang