arXiv Machine Learning

MEGA-CL: A Molecular Foundation Model for Generalizable ADMET Prediction through Graph External Attention and Contrastive Learning

arXiv:2607. 24314v1 Announce Type: new Abstract: Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery.

arXiv AI
Sep 25

SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion

SMILESGNN is a multimodal architecture that fuses a SMILES Transformer encoder with a GATv2 graph encoder through cross‑attention, enabling interpretable clinical toxicity predictions. The model retains an explicit graph branch, allowing GNNExplainer to identify substructures linked to toxicity. On the ClinTox dataset it achieves an AUC‑ROC of 0.987 and F1 of 0.906 with only 0.4 M parameters, while on Tox21 it attains a mean AUC‑ROC of 0.750, comparable to strong single‑modality baselines.

By Quang Minh Nguyen, Thuy Quynh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thanh Long Dai Doan, Trong Nghia Nguyen
arXiv Machine Learning
Jun 9

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

arXiv:2604. 26498v3 Announce Type: replace Abstract: The rapid growth of molecular foundation models and large language models (LLMs) has encouraged a scale centred view of AI in drug discovery, in which larger pretrained models are expected to supersede compact cheminformatics models.

By Jinjiang Guo, Sheng Ding
arXiv Machine Learning
Aug 6

Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs for Blood-Brain Barrier Permeability Prediction

arXiv:2608. 04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing, the blood-brain barrier.

By Marco Vieto Vega, Long D. Nguyen, Binh P. Nguyen
arXiv Machine Learning
Sep 1

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

ToxLens is a reproducible multi‑task graph‑learning framework designed for leakage‑aware, uncertainty‑calibrated prediction of 11 molecular toxicity endpoints, including Ames mutagenicity and hERG inhibition. The workflow integrates conservative chemical curation, sphere‑exclusion filtering, a leakage‑aware UMAP‑HDBSCAN split, parallel graph and global‑feature encoders with late concatenation, temperature‑scaled Monte Carlo dropout, conformal‑style prediction sets, applicability‑domain analysis, and SHAP‑guided toxicophore discovery with occlusion controls. On a leakage‑controlled test fold, a five‑seed soft‑voting ensemble achieved MCC 0.44, AUROC 0.83, and AUPRC 0.58, outperforming four ECFP4‑based shallow baselines across all endpoints.

By Magnus H. Str{\o}mme, Alex G. C. de S\'a, David B. Ascher
arXiv AI
Jul 7

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

arXiv:2607. 04557v1 Announce Type: cross Abstract: Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles.

By Dongmin Bang, Sugyun An, Inyoung Sung, Ilho Yun, Sun Kim, Sangseon Lee
arXiv Machine Learning
Jul 3

An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility

arXiv:2607. 02212v1 Announce Type: cross Abstract: Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph information into a single representation, obscuring whether a prediction is driven by global chemistry, molecular structure, or both.

By Sampreeti Bhattacharya, Arkaprava Roy