hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.
By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv:2510. 00192v3 Announce Type: replace Abstract: Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full fine-tuning.
By Xin Yu, Cong Xie, Xunmei Liu, Tiantian Fan, Lingzhou Xue, Zhi Zhang
arXiv:2601. 16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments.
By Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu
arXiv:2609.05885v1 Announce Type: new
Abstract: Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR...
By Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong
ChainDoRA is a new parameter‑efficient fine‑tuning framework for large language models that replaces the dense low‑rank factorization of LoRA with a connected Tensor‑Train (TT) chain. By separating weight magnitude and direction and using a TT rank to control representation capacity, ChainDoRA achieves higher average accuracy on seven commonsense reasoning benchmarks while dramatically reducing trainable parameters—down to 5.35 M versus 56 M for LoRA and DoRA. Ablation studies show that the TT parameterization offers controllable trade‑offs between parameter cost and accuracy.
By Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam
arXiv:2607. 01170v1 Announce Type: cross Abstract: Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces.
By Zhuoxuan Zhang (Yang), Kangqi Ni (Yang), Yuhang Chen (Yang), Mingfu Liang (Yang), Xiaohan Wei (Yang), Yunchen Pu (Yang), Fei Tian (Yang), Chonglin Sun (Yang), Frank Shyu (Yang), Adam (Yang), Song, Sandeep Pandey, Luke Simon, Tianlong Chen, Xi Liu