X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding
Read the original on arXiv Computation and Language →The paper introduces X-CoSD, a communication‑efficient framework for collaborative speculative decoding that allows a small on‑device language model to draft tokens while a server‑side large language model verifies them, even when the two models use different vocabularies. X-CoSD employs hybrid resampling to limit distribution exchange to only the common vocabulary, and its enhanced variant X-CoSD‑E further reduces communication by sending only server‑sampled replacement candidates for device verification. Both variants preserve the server model’s distribution and, according to experiments, accelerate token generation without sacrificing quality.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.