arXiv AI By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon

hLLM: Single Pass Decoding for Generative Reranking

Read the original on arXiv AI →

hLLM (Hungarian LLM) is a decoding strategy that efficiently produces ranked lists from large language models by decoding all ordinal positions in a single forward pass. It extracts an item‑position score matrix from the model’s hidden states, then applies the Hungarian algorithm to obtain a valid permutation, avoiding the sequential token generation of traditional autoregressive decoding. Experiments show that LoRA‑based fine‑tuning with teacher ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality comparable to the teacher model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

SPD: Single Pass Decoding for Generative Reranking

The paper introduces SPD (Single Forward Pass), a decoding strategy that efficiently produces ranked lists from large language models by emitting only the ordinal positions of items. SPD extracts an item‑position score matrix from the model’s hidden states, then solves a bipartite assignment problem via the Hungarian algorithm to generate a valid permutation in a single forward pass. Experiments show that LoRA‑based fine‑tuning combined with autoregressive ranking distillation achieves 28 ms end‑to‑end inference, a 64× speed‑up while preserving ranking quality.

By Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
arXiv AI
Jul 2

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

arXiv:2607. 01170v1 Announce Type: cross Abstract: Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces.

By Zhuoxuan Zhang (Yang), Kangqi Ni (Yang), Yuhang Chen (Yang), Mingfu Liang (Yang), Xiaohan Wei (Yang), Yunchen Pu (Yang), Fei Tian (Yang), Chonglin Sun (Yang), Frank Shyu (Yang), Adam (Yang), Song, Sandeep Pandey, Luke Simon, Tianlong Chen, Xi Liu
arXiv Computation and Language
Sep 23

ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning

ChainDoRA is a new parameter‑efficient fine‑tuning framework for large language models that replaces the dense low‑rank factorization of LoRA with a connected Tensor‑Train (TT) chain. By separating weight magnitude and direction and using a TT rank to control representation capacity, ChainDoRA achieves higher average accuracy on seven commonsense reasoning benchmarks while dramatically reducing trainable parameters—down to 5.35 M versus 56 M for LoRA and DoRA. Ablation studies show that the TT parameterization offers controllable trade‑offs between parameter cost and accuracy.

By Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam