arXiv:2607. 16553v1 Announce Type: new Abstract: Protein fold classification can be approached via sequence-based representations or structural descriptors, but direct comparisons between lightweight handcrafted descriptors and pretrained protein language model embeddings remain limited.
By Jianru Shen
arXiv:2606. 14217v1 Announce Type: new Abstract: Accurate prediction of protein-ligand binding affinity is essential for structure-based drug discovery.
By Peng-Fei Sun, Chuan-Xian Ren, Hong Yan
arXiv:2608. 04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing, the blood-brain barrier.
By Marco Vieto Vega, Long D. Nguyen, Binh P. Nguyen
The study investigates whether self‑supervised pretraining improves molecular graph neural networks by adapting the LeJEPA architecture to molecular graphs. While pretraining enhances learned representations and a frozen probe outperforms random initialization on tasks such as ogbg‑molhiv, it does not consistently boost finetuning performance across different data splits. Combining pretrained embeddings with 1024‑bit Morgan fingerprints yields modest gains, indicating that pretraining provides complementary information best exploited at the feature level.
By Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki
arXiv:2607. 07611v1 Announce Type: new Abstract: Background: Graph neural networks improve computational prediction of polypharmacy side effects, but standard binary cross-entropy training allocates equal capacity to well-classified and difficult examples, potentially missing clinically significant interactions.
By Faranak Hatami, Mousa Moradi
The paper introduces DPTM‑DT, a dual‑pretrained Transformer framework that integrates GROVER molecular graph embeddings, ESM protein language‑model embeddings, and CTD physicochemical descriptors for drug‑target prediction. It employs bidirectional cross‑modal attention to share drug‑target information and uses a single pair representation for continuous affinity regression, high‑affinity binary classification, and six‑level affinity classification. Experiments on Davis and KIBA datasets show that DPTM‑DT outperforms existing methods across regression, binary, and multiclass tasks, with ablation studies confirming the contributions of dual target representation, gated fusion, and cross‑modal attention.
By Ge Kong
SMILESGNN is a multimodal architecture that fuses a SMILES Transformer encoder with a GATv2 graph encoder through cross‑attention, enabling interpretable clinical toxicity predictions. The model retains an explicit graph branch, allowing GNNExplainer to identify substructures linked to toxicity. On the ClinTox dataset it achieves an AUC‑ROC of 0.987 and F1 of 0.906 with only 0.4 M parameters, while on Tox21 it attains a mean AUC‑ROC of 0.750, comparable to strong single‑modality baselines.
By Quang Minh Nguyen, Thuy Quynh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thanh Long Dai Doan, Trong Nghia Nguyen
arXiv:2609.37555v1 Announce Type: new
Abstract: Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical st...
By Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
arXiv:2407. 07357v3 Announce Type: replace Abstract: Predicting signed interactions in biological networks is crucial for understanding drug mechanisms and facilitating drug repurposing.
By Ziye Zhou, Meijie Wang, Lun Yu
arXiv:2607. 01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development.
By Wenbo Zhang
arXiv:2606. 14159v1 Announce Type: new Abstract: Protein-ligand binding affinity (PLA) prediction is critical in drug discovery.
By Shuai Li, Chuan-Xian Ren, Yuhao Li, Ziqi Huang, Yue Pan, Mingzhe Tang, Hong Yan
WEECFP-SuRGE introduces a 1024‑dimensional, parameter‑free continuous fingerprint that distributes each Morgan substructure across about thirty‑two signed positions in a single vector. The accompanying transformer architecture applies Substructure Rotary Graph‑distance Encoding (SuRGE), a RoPE‑like rotation based on molecular shortest‑path graph distance, to the fingerprint tokens. In benchmark tests, a seven‑model blend of this architecture achieves top rankings on the TDC ADMET leaderboard and outperforms classical fingerprints on most MoleculeNet regression tasks, while its tokenization scheme is shown to be near‑lossless and highly efficient for positional memory.
By Robert Epps