arXiv Machine Learning

Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks

The paper introduces a method to improve heuristic-based Bitcoin address clustering by using graph neural networks to generate contrastive embeddings. It releases a large Bitcoin transaction graph dataset, presents a learning framework that aligns embeddings with existing heuristics, and applies hierarchical clustering to refine clusters and detect suspicious merges. The approach offers a more modular and theoretically grounded way to analyze user-level activity on the blockchain.

arXiv Machine Learning
Sep 24

Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering

The paper presents a semi‑supervised learning framework for detecting illicit Bitcoin flows in Shared Send Mixers, using a large historical dataset of 163 million transactions. It demonstrates that the success of SSL depends on data quality rather than sheer volume, with high‑fidelity features such as KeyLinker address clustering and Shared Send Untangling complexity metrics achieving an F1 score of 0.84 on unlabeled data. The study also shows that common heuristics like One‑Time Change introduce noise, underscoring the importance of smarter feature engineering in blockchain forensics.

By Yekaterina Smolenkova, Nickolay Larionov, Nikolay Ivanov, Yury Yanovich
Hugging Face Trending Papers
Aug 20

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

The paper presents a method for early detection of fraudulent memecoins (rug pulls) on the Solana blockchain, using a dataset of 6.4 million tokens collected over seven months. It shows that most rug pulls occur within an hour of launch and that classic machine learning models, especially Gradient Boosting (XGBoost), can reliably predict them using only the first five minutes of trading data. Cross‑platform data fusion between PumpFun and Raydium further improves detection by reducing domain shift.

arXiv Machine Learning
Aug 19

Community Concealment from Graph Neural Networks

The paper introduces FCom‑DICE, a feature‑aware perturbation method that rewires influential edges and adjusts node features to hide a target community from graph neural network (GNN) inference. It shows that concealment effectiveness depends on boundary connectivity and feature similarity, and that FCom‑DICE outperforms structure‑only DICE on synthetic and real networks such as Facebook, Wikipedia, and Bitcoin Transactions while preserving key structural and feature properties.

By Dalyapraz Manatova, Pablo Moriano, L. Jean Camp