arXiv Machine Learning By Yue Han, Dianlin Wang

Beyond Rotations: AuroOFT for Expressive Quantized Orthogonal Fine-Tuning

Read the original on arXiv Machine Learning →

arXiv:2608. 05253v1 Announce Type: new Abstract: Quantized orthogonal fine-tuning (qoft) enables parameter-efficient adaptation of low-bit language models by learning structured activation rotations before frozen quantized weights.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.

By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang
arXiv AI
Jul 7

HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference

arXiv:2607. 04302v1 Announce Type: cross Abstract: We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax state in FP16.

By Hui Dong, Yanzhao Li, Jie Gao, Chunlu Li, Zhiyuan Zhang, Yupeng Sun, Zhenyuan Chen, Zhiqiang Zou