Hugging Face Trending Papers

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

Read the original on Hugging Face Trending Papers →

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token‑cost while preserving utility. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting the internal preference diagnostic.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 11

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token generation cost while preserving a model’s preference alignment. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting internal preference diagnostics. The method demonstrates that careful subspace selection can make large language models more token‑efficient without sacrificing utility.

By Dongfang Zhao
arXiv AI
Sep 7

SPD: Single Pass Decoding for Generative Reranking

The paper introduces SPD (Single Forward Pass), a decoding strategy that efficiently produces ranked lists from large language models by emitting only the ordinal positions of items. SPD extracts an item‑position score matrix from the model’s hidden states, then solves a bipartite assignment problem via the Hungarian algorithm to generate a valid permutation in a single forward pass. Experiments show that LoRA‑based fine‑tuning combined with autoregressive ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv AI
Jul 24

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

arXiv:2605. 27390v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs.

By Shuyu Zhang, Lingfeng Pan, Qicheng Wang, Yaqi Shi, Yueyang Tan, Ruyu Yan, Jiaqi Chen, Lixing Du, Lu Wang
arXiv AI
Sep 3

hLLM: Single Pass Decoding for Generative Reranking

hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon