arXiv Machine Learning

Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth

arXiv:2608. 16010v1 Announce Type: new Abstract: Model compression is critical for deploying networks on resource-constrained edge devices.

arXiv AI
Jun 9

Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

arXiv:2605. 15491v2 Announce Type: replace-cross Abstract: Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the next surviving layer and the distribution it was trained to process, leading to significant performance degradation.

By Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy, Sunwoo Lee
arXiv Machine Learning
Sep 7

From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy

The paper proposes a new strategy for merging layers in deep neural networks, enabling depth compression without requiring an analytical solution for convolutions with padding and without increasing kernel size. This approach addresses limitations of previous methods that struggled with padded convolutions and larger kernels, and it is validated across various architectures and datasets with measured inference speed-ups on embedded platforms.

By Petro Shulzhenko, Gabriele Spadaro, Enzo Tartaglione