Hugging Face Trending Papers

Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT

Read the original on Hugging Face Trending Papers →

Deploying large language models (LLMs) on Industrial Internet of Things (IIoT) edge devices demands extreme compression, yet existing structured pruning methods collapse at high compression ratios due to one-shot importance estimation, and their cross-architecture behavior remains unpredictable. This article presents a cascaded multi-granularity pruning framework that removes layers, attention heads, and feed-forward channels in coarse-to-fine order, with lightweight low-rank recovery between stages to re-estimate component importance.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.