arXiv:2609.37555v1 Announce Type: new
Abstract: Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical st...
By Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
The paper introduces NSA-Bench, a public benchmark for predicting nano self‑assembly (NSA) between molecular pairs, framing it as a binary classification problem. It presents NSA‑Net, a multimodal learning framework that fuses graph topology, sequence semantics, and physicochemical descriptors to predict self‑assembly, achieving high ROC‑AUC scores and outperforming existing baselines. The study also demonstrates how NSA‑Net’s predictions can guide experimental formulation refinement through an NSA‑Agent case study.
By Quan Hao, Mengyue Fan, Zifan Dong, Jianduo Zhao, Changhao Xiao, Shangqing Jiao, Hao Zhang, Yudong Wang, Fei Xia, Jigang Wang, Liguo Zhang, Chong Qiu
The study presents an interpretable machine‑learning framework that predicts whether lipid nanoparticles (LNPs) accumulate in the liver or in extrahepatic tissues after intravenous injection. Using a curated dataset of 476 LNP formulations, the authors engineered 808 features from lipid chemistry and formulation composition, and achieved ROC‑AUC scores up to 0.874 with tree‑based models. SHAP analysis identified ionizable‑lipid descriptors and formulation fractions—especially ionizable lipid, sterol, and PEGylated/polymer‑conjugated lipid components—as key drivers of biodistribution, offering actionable design principles for targeting tissues beyond the liver.
By Asal Mehradfar, Mohammad Shahab Sepehri, Owen Antholine, Varun Shankar, Glen S. Kwon, Salman Avestimehr, Morteza Rasoulianboroujeni
arXiv:2504. 13853v2 Announce Type: replace-cross Abstract: Rational design of lipid nanoparticles (LNPs) for tissue-specific delivery critically depends on predicting the composition of the protein corona that forms on the lipid surface after intravenous administration.
By Pingfei Zhu, Hongyi Liu, Xueyan Liu, Zhenjun Yang, Bo Yang
arXiv:2608. 11444v1 Announce Type: cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs.
By Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens
arXiv:2606. 30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets.
By Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart, Fabian Pauly
arXiv:2509. 26405v2 Announce Type: replace Abstract: We introduce InVirtuoGen, a discrete flow generative model for fragmented SMILES for de novo and fragment-constrained generation, and target-property/lead optimization of small molecules.
By Benno Kaech, Luis Wyss, Karsten Borgwardt, Gianvito Grasso
arXiv:2506. 13196v5 Announce Type: replace Abstract: Accurate prediction of protein-ligand binding affinity is critical for drug discovery.
By Han Liu, Keyan Ding, Peilin Chen, Yinwei Wei, Liqiang Nie, Dapeng Wu, Shiqi Wang
arXiv:2606. 09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once.
By Sujoy Banik, Sayantan Chakraborty, Boishakhi Das Toma, Zainab Ghafoor, Ushashi Bhattacharjee, Koushik Howlader, Tirtho Roy
SMILESGNN is a multimodal architecture that fuses a SMILES Transformer encoder with a GATv2 graph encoder through cross‑attention, enabling interpretable clinical toxicity predictions. The model retains an explicit graph branch, allowing GNNExplainer to identify substructures linked to toxicity. On the ClinTox dataset it achieves an AUC‑ROC of 0.987 and F1 of 0.906 with only 0.4 M parameters, while on Tox21 it attains a mean AUC‑ROC of 0.750, comparable to strong single‑modality baselines.
By Quang Minh Nguyen, Thuy Quynh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thanh Long Dai Doan, Trong Nghia Nguyen
arXiv:2607. 00464v1 Announce Type: new Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules.
By Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding, Huajun Chen
arXiv:2608. 13797v1 Announce Type: new Abstract: Computational approaches to drug discovery involve multiple sub-problems, and among them, drug-target binding affinity prediction plays an important role.
By Jafin Khan, Md Hossain Shuvo