arXiv Machine Learning

Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics

arXiv:2608. 12864v1 Announce Type: cross Abstract: Public blockchain data enables large-scale DeFi-related analysis, but many existing approaches are application-specific, difficult to scale, or hard to interpret.

Hugging Face Trending Papers
Aug 20

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

The paper presents a method for early detection of fraudulent memecoins (rug pulls) on the Solana blockchain, using a dataset of 6.4 million tokens collected over seven months. It shows that most rug pulls occur within an hour of launch and that classic machine learning models, especially Gradient Boosting (XGBoost), can reliably predict them using only the first five minutes of trading data. Cross‑platform data fusion between PumpFun and Raydium further improves detection by reducing domain shift.

arXiv Machine Learning
Sep 24

Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering

The paper presents a semi‑supervised learning framework for detecting illicit Bitcoin flows in Shared Send Mixers, using a large historical dataset of 163 million transactions. It demonstrates that the success of SSL depends on data quality rather than sheer volume, with high‑fidelity features such as KeyLinker address clustering and Shared Send Untangling complexity metrics achieving an F1 score of 0.84 on unlabeled data. The study also shows that common heuristics like One‑Time Change introduce noise, underscoring the importance of smarter feature engineering in blockchain forensics.

By Yekaterina Smolenkova, Nickolay Larionov, Nikolay Ivanov, Yury Yanovich
arXiv AI
Sep 12

DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks

DeFiFusion is a dual‑modal framework that detects price manipulation attacks in decentralized finance by jointly analyzing transaction events and smart contract semantics. It encodes fine‑grained temporal and economic features from transactions and extracts contract logic using large language models, then fuses these signals with a Dual‑Modal Projection‑Fusion Transformer. The approach achieves state‑of‑the‑art performance, recalling 222 of 225 known attacks with 96.10% precision.

By Rui Cao, Shaojing Fan, Liming Fang, Yuchan Liu, Yingying Jiao, Zhenguang Liu
arXiv Computation and Language
Sep 1

TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding

TxSum introduces a user-centered approach to understanding Ethereum transactions by providing structured, risk-aware explanations grounded at the token‑flow level. The authors built a dataset of 187 complex transactions with 2,375 token‑flow annotations and transaction‑level summaries, and developed MATEX, a multi‑agent framework that retrieves external knowledge and audits explanations for factual consistency. MATEX outperforms existing baselines, improving user comprehension from 52.9% to 76.5% and increasing malicious‑transaction rejection from 36.0% to 88.0% while keeping false‑rejection rates low.

By Zifan Peng, Jingyi Zheng, Yule Liu, Huaiyu Jia, Qiming Ye, Jingyu Liu, Xufeng Yang, Mingchen Li, Qingyuan Gong, Xuechao Wang, Xinlei He
arXiv AI
Jul 9

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607. 06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks.

By Kiarash Ahi, Saeed Valizadeh