Hugging Face Trending Papers

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token‑cost while preserving utility. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting the internal preference diagnostic.

arXiv Machine Learning
Sep 11

LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

LOCUS is a post‑training technique that selects a task‑aware low‑rank adaptation subspace to reduce token generation cost while preserving a model’s preference alignment. By updating only a tiny fraction of parameters (0.24–0.28 %) on 3 B‑parameter backbones, LOCUS cuts continuation length by up to 39.84 % on Pythia‑2.8B and 14.87–17.58 % on Qwen2.5‑3B, without materially affecting internal preference diagnostics. The method demonstrates that careful subspace selection can make large language models more token‑efficient without sacrificing utility.

By Dongfang Zhao
arXiv AI
Sep 7

SPD: Single Pass Decoding for Generative Reranking

The paper introduces SPD (Single Forward Pass), a decoding strategy that efficiently produces ranked lists from large language models by emitting only the ordinal positions of items. SPD extracts an item‑position score matrix from the model’s hidden states, then solves a bipartite assignment problem via the Hungarian algorithm to generate a valid permutation in a single forward pass. Experiments show that LoRA‑based fine‑tuning combined with autoregressive ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv AI
Jul 24

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

arXiv:2605. 27390v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs.

By Shuyu Zhang, Lingfeng Pan, Qicheng Wang, Yaqi Shi, Yueyang Tan, Ruyu Yan, Jiaqi Chen, Lixing Du, Lu Wang
arXiv AI
Sep 3

hLLM: Single Pass Decoding for Generative Reranking

hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv AI
Jun 11

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

arXiv:2605. 12288v3 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences over full sequences even though generation is driven by per-token decisions.

By Truong Nguyen, Tien-Phat Nguyen, Linh Ngo Van, Duy Minh Ho Nguyen, Khoa Doan, Trung Le
arXiv AI
2d ago

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

DRelay introduces a global draft context mechanism to improve prefix-aware parallel speculative decoding for large language models. By using a global reader to extract predictive information across the entire draft block and a causal selector to repair early token selection errors, DRelay extends the accepted prefix length and enhances decoding performance. Experiments on eight benchmarks show consistent gains over existing methods such as DFlash, Domino, and DSpark, with notable speedup improvements in SGLang serving.

By Zhuoyu Wang, Junnan Huang, Xinyu Chen