arXiv Machine Learning

An Introduction to Compression-Based Machine Learning

The paper discusses how any lossless compression algorithm can be transformed into a machine learning method using Normalized Compression Distance or the Minimum Description Length principle, and conversely how any auto‑regressive model can become a lossless compressor via entropy coding. It surveys and formalizes these strategies, introduces a design framework for compression‑based ML, and empirically validates that such methods can match conventional baselines and outperform them on malware detection, achieving accuracy gains up to 0.62 by varying design choices.

arXiv Machine Learning
Sep 24

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

The paper presents a lightweight machine‑learning approach for multi‑class malware detection on resource‑constrained devices. Using a LightGBM classifier with SMOTE oversampling, SOM‑US undersampling, and Genetic‑Algorithm feature selection, the authors achieve 89.1 % accuracy on four malware families and 76 % on 16 individual malware types. A second Random‑Forest model further improves family classification to 91.2 % and individual classification to 78.7 %.

By Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain
arXiv AI
4d ago

HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

HyperZip introduces an efficient data compression framework that uses diffusion-based large language models (dLLMs) with Multi-Token Prediction to speed up compression. It addresses the trade‑off between throughput and compression rate by employing a hypernetwork that generates data‑specific updates from a context representation, allowing the dLLM to adapt to target data without costly fine‑tuning. Experiments show HyperZip outperforms state‑of‑the‑art baselines in both compression rate and speed.

By Thai Nguyen, Khang Tran, NhatHai Phan
arXiv AI
Aug 13

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

arXiv:2608. 11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression.

By Angelo Nardone, Paolo Ferragina
arXiv Machine Learning
Sep 18

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.

By Bijied Brahimi, Vincent Cohadon, Gabriel Glazman, Rayan Al Mohaize, Omran Berjawi, Rida Khatoun