arXiv Machine Learning By Robert Epps

WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding

Read the original on arXiv Machine Learning →

WEECFP-SuRGE introduces a 1024‑dimensional, parameter‑free continuous fingerprint that distributes each Morgan substructure across about thirty‑two signed positions in a single vector. The accompanying transformer architecture applies Substructure Rotary Graph‑distance Encoding (SuRGE), a RoPE‑like rotation based on molecular shortest‑path graph distance, to the fingerprint tokens. In benchmark tests, a seven‑model blend of this architecture achieves top rankings on the TDC ADMET leaderboard and outperforms classical fingerprints on most MoleculeNet regression tasks, while its tokenization scheme is shown to be near‑lossless and highly efficient for positional memory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
3d ago

WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction

WEECFP-SuRGE introduces a position‑aware substructure encoding method that combines tokenized hierarchical Morgan fingerprints with graph‑distance‑dependent rotations applied at the input and within transformer self‑attention. The approach captures local chemistry, long‑range interactions, and molecular topology without requiring external pretraining or 3‑D conformer generation. Benchmarks on MoleculeNet and the Therapeutic Data Commons ADMET datasets show competitive performance, and a reconstruction procedure correctly identifies constitutional isomers for 92.6% of a 4,200‑molecule library.

By Robert Epps
arXiv Machine Learning
Sep 7

Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA

The study investigates whether self‑supervised pretraining improves molecular graph neural networks by adapting the LeJEPA architecture to molecular graphs. While pretraining enhances learned representations and a frozen probe outperforms random initialization on tasks such as ogbg‑molhiv, it does not consistently boost finetuning performance across different data splits. Combining pretrained embeddings with 1024‑bit Morgan fingerprints yields modest gains, indicating that pretraining provides complementary information best exploited at the feature level.

By Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer