arXiv Machine Learning

A Machine Learning Benchmarking Framework for Lipid Nanoparticle Transfection Efficiency Prediction

arXiv:2507. 03209v2 Announce Type: replace-cross Abstract: The discovery of new ionizable lipids for efficient lipid nanoparticle (LNP)-mediated RNA delivery remains a major bottleneck in RNA therapeutics development.

arXiv Statistics ML
Sep 7

Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction

The paper introduces NSA-Bench, a public benchmark for predicting nano self‑assembly (NSA) between molecular pairs, framing it as a binary classification problem. It presents NSA‑Net, a multimodal learning framework that fuses graph topology, sequence semantics, and physicochemical descriptors to predict self‑assembly, achieving high ROC‑AUC scores and outperforming existing baselines. The study also demonstrates how NSA‑Net’s predictions can guide experimental formulation refinement through an NSA‑Agent case study.

By Quan Hao, Mengyue Fan, Zifan Dong, Jianduo Zhao, Changhao Xiao, Shangqing Jiao, Hao Zhang, Yudong Wang, Fei Xia, Jigang Wang, Liguo Zhang, Chong Qiu
arXiv Machine Learning
Sep 17

Decoding Extrahepatic Targeting of Lipid Nanoparticles with Interpretable Machine Learning

The study presents an interpretable machine‑learning framework that predicts whether lipid nanoparticles (LNPs) accumulate in the liver or in extrahepatic tissues after intravenous injection. Using a curated dataset of 476 LNP formulations, the authors engineered 808 features from lipid chemistry and formulation composition, and achieved ROC‑AUC scores up to 0.874 with tree‑based models. SHAP analysis identified ionizable‑lipid descriptors and formulation fractions—especially ionizable lipid, sterol, and PEGylated/polymer‑conjugated lipid components—as key drivers of biodistribution, offering actionable design principles for targeting tissues beyond the liver.

By Asal Mehradfar, Mohammad Shahab Sepehri, Owen Antholine, Varun Shankar, Glen S. Kwon, Salman Avestimehr, Morteza Rasoulianboroujeni
arXiv AI
Sep 25

SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion

SMILESGNN is a multimodal architecture that fuses a SMILES Transformer encoder with a GATv2 graph encoder through cross‑attention, enabling interpretable clinical toxicity predictions. The model retains an explicit graph branch, allowing GNNExplainer to identify substructures linked to toxicity. On the ClinTox dataset it achieves an AUC‑ROC of 0.987 and F1 of 0.906 with only 0.4 M parameters, while on Tox21 it attains a mean AUC‑ROC of 0.750, comparable to strong single‑modality baselines.

By Quang Minh Nguyen, Thuy Quynh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thanh Long Dai Doan, Trong Nghia Nguyen