arXiv Statistics ML By Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara

IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs

Read the original on arXiv Statistics ML →

IrekoGPT is a post‑hoc technique that transforms pretrained large language models into slimmable versions, enabling dynamic width adjustment during inference. It builds on SliceGPT by keeping the original projection matrices intact, thereby exposing nested subnetworks at various widths. The method enhances robustness through layer‑wise calibration across multiple compression ratios and refines downstream linear layers using gradient‑free ridge regression, yielding better performance than naive PCA‑based slimming on Llama and Qwen models, especially at high compression levels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.

arXiv Machine Learning
Jul 28

Compressing LLMs with MoP: Mixture of Pruners

arXiv:2602. 06127v2 Announce Type: replace Abstract: The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference.

By Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Leandro Giusti Mugnaini, Keith Ando Ogawa, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao