arXiv Computation and Language By David Ponce, Thierry Etchegoyhen, Javier Del Ser

GeLaCo: An Evolutionary Approach to Layer Compression

Read the original on arXiv Computation and Language →

GeLaCo is an evolutionary method for compressing large language models by collapsing layers through parametrized weight merging. It uses population-based search with a fitness function that balances similarity of residual updates and language modeling KL divergence, enabling both single and multi-objective compression. The approach yields Pareto-optimal trade-offs between compression and quality, outperforming existing methods in perplexity and generative evaluations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 21

NIRVANA: Structured Pruning Reimagined for Large Language Model Compression

arXiv:2509. 14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retraining to recover capabilities.

By Mengting Ai, Tianxin Wei, Sirui Chen, Jingrui He
arXiv Computation and Language
Aug 31

Pruning Laws for Large Language Models

arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly gro...

By Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty