arXiv AI

Pruned Traffic Trees: Native Semantic Compression with a Protocol-Structured Model Family for Encrypted Traffic Classification

arXiv Machine Learning
Aug 6

Learning Compression Rules for Network Traffic

arXiv:2608. 04545v1 Announce Type: new Abstract: We study the problem of learning compact rule-based compressors for structured network traffic.

By Quentin Lampin (Orange Research), \'Eloi Sainte-Beuve (Orange Research, Universit\'e Grenoble Alpes), Louis-Adrien Dufr\`ene (Orange Research), Guillaume Larue (Orange Research), Massih-Reza Amini (Universit\'e Grenoble Alpes)
Hugging Face Trending Papers
Jun 25

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization methods that rely on data-intensive and costly quantization-aware training to mitigate severe performance degradation, CAT-Q is a simple yet effective post-training quantization scheme that is readily applicable to LLMs with diverse architectures and model sizes.

arXiv AI
Jun 10

Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

arXiv:2606. 09924v1 Announce Type: cross Abstract: Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather than computation: the dense network cannot be retained on-chip, and every parameter must be loaded for every input.

By Kohga Tanaka, Hiroaki Nishi