arXiv Machine Learning

WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding

WEECFP-SuRGE introduces a 1024‑dimensional, parameter‑free continuous fingerprint that distributes each Morgan substructure across about thirty‑two signed positions in a single vector. The accompanying transformer architecture applies Substructure Rotary Graph‑distance Encoding (SuRGE), a RoPE‑like rotation based on molecular shortest‑path graph distance, to the fingerprint tokens. In benchmark tests, a seven‑model blend of this architecture achieves top rankings on the TDC ADMET leaderboard and outperforms classical fingerprints on most MoleculeNet regression tasks, while its tokenization scheme is shown to be near‑lossless and highly efficient for positional memory.

arXiv Machine Learning
3d ago

WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction

WEECFP-SuRGE introduces a position‑aware substructure encoding method that combines tokenized hierarchical Morgan fingerprints with graph‑distance‑dependent rotations applied at the input and within transformer self‑attention. The approach captures local chemistry, long‑range interactions, and molecular topology without requiring external pretraining or 3‑D conformer generation. Benchmarks on MoleculeNet and the Therapeutic Data Commons ADMET datasets show competitive performance, and a reconstruction procedure correctly identifies constitutional isomers for 92.6% of a 4,200‑molecule library.

By Robert Epps
arXiv Machine Learning
Sep 7

Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA

The study investigates whether self‑supervised pretraining improves molecular graph neural networks by adapting the LeJEPA architecture to molecular graphs. While pretraining enhances learned representations and a frozen probe outperforms random initialization on tasks such as ogbg‑molhiv, it does not consistently boost finetuning performance across different data splits. Combining pretrained embeddings with 1024‑bit Morgan fingerprints yields modest gains, indicating that pretraining provides complementary information best exploited at the feature level.

By Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
arXiv Machine Learning
Sep 22

CurvFlow-DTA: dual-graph discrete Ricci curvature flow for drug--target affinity prediction

CurvFlow-DTA introduces a dual-graph discrete Ricci curvature flow framework for drug–target affinity prediction, replacing static curvature with weighted Forman curvature flow on both drug and protein residue–residue contact graphs. The method precomputes a label‑independent flow trajectory for each entity and uses a pair‑conditioned selector to guide a dual‑branch Flow‑GINE, leveraging frozen ESM‑2 residue representations. Experiments on Davis and KIBA datasets show significant improvements over the Ricci‑GraphDTA baseline, with reductions in mean squared error of up to 19.9% in warm‑start and 27.4% in cold‑start settings, and higher concordance indices across benchmarks.

By Jicheng Ma, Yunyan Yang, Juan Zhao, Liang Zhao
arXiv AI
Jul 29

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.

By Kai Lun Huang (California State University, Fullerton), Wei Chieh Sun (University of Washington)
arXiv Machine Learning
Aug 6

Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs for Blood-Brain Barrier Permeability Prediction

arXiv:2608. 04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing, the blood-brain barrier.

By Marco Vieto Vega, Long D. Nguyen, Binh P. Nguyen