LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token‑cost while preserving utility. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting the internal preference diagnostic.
arXiv:2605. 27390v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs.
By Shuyu Zhang, Lingfeng Pan, Qicheng Wang, Yaqi Shi, Yueyang Tan, Ruyu Yan, Jiaqi Chen, Lixing Du, Lu Wang
The paper introduces SPD (Single Forward Pass), a decoding strategy that efficiently produces ranked lists from large language models by emitting only the ordinal positions of items. SPD extracts an item‑position score matrix from the model’s hidden states, then solves a bipartite assignment problem via the Hungarian algorithm to generate a valid permutation in a single forward pass. Experiments show that LoRA‑based fine‑tuning combined with autoregressive ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality.
By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv:2509. 23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences.
By Lucio La Cava, Andrea Tagarelli
hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.
By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv:2602. 09492v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical gains, often on the same benchmarks.
By Sangyoon Lee, Jaeho Lee