arXiv Machine Learning

Decoding Extrahepatic Targeting of Lipid Nanoparticles with Interpretable Machine Learning

The study presents an interpretable machine‑learning framework that predicts whether lipid nanoparticles (LNPs) accumulate in the liver or in extrahepatic tissues after intravenous injection. Using a curated dataset of 476 LNP formulations, the authors engineered 808 features from lipid chemistry and formulation composition, and achieved ROC‑AUC scores up to 0.874 with tree‑based models. SHAP analysis identified ionizable‑lipid descriptors and formulation fractions—especially ionizable lipid, sterol, and PEGylated/polymer‑conjugated lipid components—as key drivers of biodistribution, offering actionable design principles for targeting tissues beyond the liver.

arXiv Machine Learning
Jul 17

A Machine Learning Benchmarking Framework for Lipid Nanoparticle Transfection Efficiency Prediction

arXiv:2507. 03209v2 Announce Type: replace-cross Abstract: The discovery of new ionizable lipids for efficient lipid nanoparticle (LNP)-mediated RNA delivery remains a major bottleneck in RNA therapeutics development.

By Asal Mehradfar, Mohammad Shahab Sepehri, Jose Miguel Hernandez-Lobato, Glen S. Kwon, Mahdi Soltanolkotabi, Salman Avestimehr, Morteza Rasoulianboroujeni
arXiv Machine Learning
Jun 5

A differentiable machine learning small-angle X-ray scattering analysis framework for structure elucidation of lipid nanoparticles

arXiv:2606. 05200v1 Announce Type: cross Abstract: Lipid nanoparticles (LNPs) are efficient delivery systems for negatively charged nucleic acids.

By Maria B{\aa}nkestad, Sandra Barman, Magnus R\"oding, Erik Kaunisto, Viktoriia Meklesh, Audrey Gallud, Marco Mendez, Marianna Yanez Arteta, Stefan Norberg, Ann Terry, Smita Chakraborty, Shun Yu, Jerk R\"onnols, Sepideh Pashami
arXiv Machine Learning
Aug 7

Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles

arXiv:2608. 05761v1 Announce Type: new Abstract: The development of nanotherapeutics often involves extensive empirical optimization due to the sensitivity of nanoparticle properties, such as size and polydispersity index (PDI), to minor changes in process parameters.

By Kai Dahms, Eilien Heinrich, Jochen Schmid, Michael Bortz, Iryna Savych, Regina Bleul
arXiv Machine Learning
Sep 1

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

ToxLens is a reproducible multi‑task graph‑learning framework designed for leakage‑aware, uncertainty‑calibrated prediction of 11 molecular toxicity endpoints, including Ames mutagenicity and hERG inhibition. The workflow integrates conservative chemical curation, sphere‑exclusion filtering, a leakage‑aware UMAP‑HDBSCAN split, parallel graph and global‑feature encoders with late concatenation, temperature‑scaled Monte Carlo dropout, conformal‑style prediction sets, applicability‑domain analysis, and SHAP‑guided toxicophore discovery with occlusion controls. On a leakage‑controlled test fold, a five‑seed soft‑voting ensemble achieved MCC 0.44, AUROC 0.83, and AUPRC 0.58, outperforming four ECFP4‑based shallow baselines across all endpoints.

By Magnus H. Str{\o}mme, Alex G. C. de S\'a, David B. Ascher
arXiv Machine Learning
Jul 15

Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

arXiv:2607. 12349v1 Announce Type: new Abstract: Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design.

By Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen, Frazier N. Baker, David C. Kombo, John L. Kane Jr., Andrew A. Scholte, Yi Li, Matthew J. LaMarche, Luigi I. Iconaru, Hans-Peter Biemann, Mingyi Hong, Xia Ning
arXiv Statistics ML
Sep 7

Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction

The paper introduces NSA-Bench, a public benchmark for predicting nano self‑assembly (NSA) between molecular pairs, framing it as a binary classification problem. It presents NSA‑Net, a multimodal learning framework that fuses graph topology, sequence semantics, and physicochemical descriptors to predict self‑assembly, achieving high ROC‑AUC scores and outperforming existing baselines. The study also demonstrates how NSA‑Net’s predictions can guide experimental formulation refinement through an NSA‑Agent case study.

By Quan Hao, Mengyue Fan, Zifan Dong, Jianduo Zhao, Changhao Xiao, Shangqing Jiao, Hao Zhang, Yudong Wang, Fei Xia, Jigang Wang, Liguo Zhang, Chong Qiu
arXiv AI
Sep 3

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI is a new framework for drug‑target interaction prediction that uses probe‑driven pattern matching to preserve weak biochemical signals. It introduces IterProbe, which retains contextual states across refinement depths and selects them with learnable probes, and BindingProbe, which models drug‑protein complementarity at both local and whole‑pair levels. Experiments show that ProbeMatchDTI outperforms existing methods, improving AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank, and its predictions can be integrated into downstream drug‑discovery workflows.

By Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
arXiv Machine Learning
Sep 25

Feature Space Selection and Heterogeneous Effect Estimation for Blood-Brain Barrier Permeability: A Random Forest to the Generalized Random Forest Pipeline

The study evaluates how different molecular feature spaces—Morgan fingerprints, RDKit physicochemical descriptors, and SMILES bigrams—affect the prediction of blood‑brain barrier permeability using various learning algorithms. Dynamic Random Forests with combined features achieved the best performance (mean AUC 0.970). When applying Generalized Random Forests to estimate heterogeneous effects of LogP on BBB permeability, orthogonalization revealed that apparent heterogeneity largely vanished after accounting for confounding, shifting importance toward residual SMILES bigram information.

By Tshemollo Rapolai, Seite Makgai, Mohammad Arashi
Hugging Face Trending Papers
Sep 2

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.