IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs
Read the original on arXiv Statistics ML →IrekoGPT is a post‑hoc technique that transforms pretrained large language models into slimmable versions, enabling dynamic width adjustment during inference. It builds on SliceGPT by keeping the original projection matrices intact, thereby exposing nested subnetworks at various widths. The method enhances robustness through layer‑wise calibration across multiple compression ratios and refines downstream linear layers using gradient‑free ridge regression, yielding better performance than naive PCA‑based slimming on Llama and Qwen models, especially at high compression levels.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.