Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task accuracy does each efficiency knob actually cost?
Post-training quantization lowers the memory footprint of Large Language Models (LLMs) and speeds up inference, which is why it is now common for on-device deployment. Most of what we know about its e...
arXiv:2609.00224v1 Announce Type: cross
Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existi...
By Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi
The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.
By Anirudh Malik, M Sparsh Mehra, Poojith Devan
The paper reports a post‑training ternarization of the 4‑billion‑parameter Qwen model, achieving an effective 1.641‑bit representation for 81.62 % of its weights while keeping activations at 16‑bit precision. Accuracy drops from 64.5 % to 54.7 % across ten capability tests, with uneven degradation (e.g., BoolQ 84.6 % of teacher performance, ARC‑Challenge 43.8 %). After packing the ternary planes, the model size shrinks from 8.29 GiB to 3.96 GiB with negligible change in perplexity, though inference speed is not improved.
By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv:2605.07111v3 Announce Type: replace-cross
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
By Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li, Virginia Smith, Kevin Kuo
The paper presents a Pareto atlas of LLM inference optimizations, mapping cost, quality, and latency trade‑offs for Qwen2.5‑7B‑Instruct on L4, A100, and H100 GPUs. Using 54 measured configurations and a calibrated simulator, it identifies 18 of 36 setups on the Pareto frontier, showing that combined methods outperform single ones. Quality tests reveal that AWQ 4bit and FP8 weights offer significant latency reductions while largely preserving accuracy, but naive FP8 KV caching fails to answer any questions correctly.
By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv:2609.39223v2 Announce Type: new
Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...
By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.
By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
Task-Aware Spectral Pruning (TASP) is a post‑training framework that tailors sparse masks to specific tasks by calibrating module‑level spectral descriptors against task‑specific ablation effects. It constructs masks that close grouped‑query‑attention and SwiGLU dependencies, routing each user turn to a single compiled mask that remains fixed during prefill and decoding. In experiments, TASP achieves a 43% active‑FLOP reduction while preserving 97.7% of the dense BF16 performance on Llama‑3‑70B, and delivers a 1.44× speedup on an A100 80GB with INT8‑weight/BF16‑compute, reducing decode latency from 45.2 to 31.3 ms/token.
By Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma
arXiv:2608.28911v1 Announce Type: new
Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length....
By Daeha Lee, Do-Hyung Kim, Jae-Hong Kim
arXiv:2606. 13767v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models.
By Elijah Cadenhead, Cristian McGee, Xin Li, El Houcine Bergou, Aritra Dutta