arXiv Machine Learning By John Hurwitz, Edward Raff, Charles K. Nicholas

An Introduction to Compression-Based Machine Learning

Read the original on arXiv Machine Learning →

The paper discusses how any lossless compression algorithm can be transformed into a machine learning method using Normalized Compression Distance or the Minimum Description Length principle, and conversely how any auto‑regressive model can become a lossless compressor via entropy coding. It surveys and formalizes these strategies, introduces a design framework for compression‑based ML, and empirically validates that such methods can match conventional baselines and outperform them on malware detection, achieving accuracy gains up to 0.62 by varying design choices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv AI
4d ago

HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

HyperZip introduces an efficient data compression framework that uses diffusion-based large language models (dLLMs) with Multi-Token Prediction to speed up compression. It addresses the trade‑off between throughput and compression rate by employing a hypernetwork that generates data‑specific updates from a context representation, allowing the dLLM to adapt to target data without costly fine‑tuning. Experiments show HyperZip outperforms state‑of‑the‑art baselines in both compression rate and speed.

By Thai Nguyen, Khang Tran, NhatHai Phan