arXiv Machine Learning By Dongfang Zhao

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

Read the original on arXiv Machine Learning →

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token generation cost while preserving a model’s preference alignment. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting internal preference diagnostics. The method demonstrates that careful subspace selection can make large language models more token‑efficient without sacrificing utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 10

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token‑cost while preserving utility. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting the internal preference diagnostic.

arXiv AI
Jul 24

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

arXiv:2605. 27390v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs.

By Shuyu Zhang, Lingfeng Pan, Qicheng Wang, Yaqi Shi, Yueyang Tan, Ruyu Yan, Jiaqi Chen, Lixing Du, Lu Wang
arXiv AI
Sep 7

SPD: Single Pass Decoding for Generative Reranking

The paper introduces SPD (Single Forward Pass), a decoding strategy that efficiently produces ranked lists from large language models by emitting only the ordinal positions of items. SPD extracts an item‑position score matrix from the model’s hidden states, then solves a bipartite assignment problem via the Hungarian algorithm to generate a valid permutation in a single forward pass. Experiments show that LoRA‑based fine‑tuning combined with autoregressive ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv AI
Sep 3

hLLM: Single Pass Decoding for Generative Reranking

hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon