arXiv Machine Learning

Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering

The paper presents a semi‑supervised learning framework for detecting illicit Bitcoin flows in Shared Send Mixers, using a large historical dataset of 163 million transactions. It demonstrates that the success of SSL depends on data quality rather than sheer volume, with high‑fidelity features such as KeyLinker address clustering and Shared Send Untangling complexity metrics achieving an F1 score of 0.84 on unlabeled data. The study also shows that common heuristics like One‑Time Change introduce noise, underscoring the importance of smarter feature engineering in blockchain forensics.

Hugging Face Trending Papers
Aug 20

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

The paper presents a method for early detection of fraudulent memecoins (rug pulls) on the Solana blockchain, using a dataset of 6.4 million tokens collected over seven months. It shows that most rug pulls occur within an hour of launch and that classic machine learning models, especially Gradient Boosting (XGBoost), can reliably predict them using only the first five minutes of trading data. Cross‑platform data fusion between PumpFun and Raydium further improves detection by reducing domain shift.

arXiv Machine Learning
Sep 3

Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks

The paper introduces a method to improve heuristic-based Bitcoin address clustering by using graph neural networks to generate contrastive embeddings. It releases a large Bitcoin transaction graph dataset, presents a learning framework that aligns embeddings with existing heuristics, and applies hierarchical clustering to refine clusters and detect suspicious merges. The approach offers a more modular and theoretically grounded way to analyze user-level activity on the blockchain.

By Hugo Schnoering, Roman Bresson, Michalis Vazirgiannis
arXiv Machine Learning
Jun 29

CO-DEFEND: Continuous Decentralized Federated Learning for Secure DoH-Based Threat Detection

arXiv:2504. 01882v2 Announce Type: replace Abstract: The use of DNS over HTTPS (DoH) tunneling by an attacker to hide malicious activity within encrypted DNS traffic poses a serious threat to network security, as it allows malicious actors to bypass traditional monitoring and intrusion detection systems while evading detection by conventional traffic analysis techniques.

By Diego Cajaraville-Aboy, Marta Moure-Garrido, Carlos Beis-Penedo, Carlos Garcia-Rubio, Rebeca P. D\'iaz-Redondo, Celeste Campo, Ana Fern\'andez-Vilas, Manuel Fern\'andez-Veiga
arXiv Machine Learning
Sep 15

End-to-End Verifiable and Robust Federated Learning

arXiv:2609.15521v1 Announce Type: new Abstract: Federated learning enables multiple parties to train a shared model without centralizing raw data with the help of an aggregator, but introduces integr...

By Doryan Lesaignoux, Enrique M\'armol Campos, Gabriele Spini, Jos\'e L. Hern\'andez-Ramos, Stephan Krenn
arXiv Machine Learning
Jul 27

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

arXiv:2607. 21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data.

By Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot