arXiv Machine Learning By Miroslav L\v{z}i\v{c}a\v{r} (Deep MedChem)

OPDiv: Optimal Selection of Top-K High-Scoring, Diverse Compounds

Read the original on arXiv Machine Learning →

OPDiv is a new algorithm that addresses the tradeoff between ranking quality and chemical diversity in virtual screening. It uses integer optimization to select an optimal subset of top‑k compounds that meet a specified diversity threshold, evaluated with fingerprint, shape, and electrostatic metrics. The method demonstrates how to benchmark virtual screening pipelines by comparing their best achievable diverse selections.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 26

Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical Space

arXiv:2606. 26657v1 Announce Type: new Abstract: Identifying high-utility candidates from massive discrete spaces under expensive evaluations is a recurring challenge across the sciences, with structure-based drug discovery as a prominent example.

By Mohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay, Joana Reis, Felix Strieth-Kalthoff, Haribabu Arthanari
arXiv Machine Learning
Sep 7

Small Molecule Optimization with Large Language Models

The paper introduces Mol-E, an evolutionary algorithm that leverages large language models trained on molecular data to generate candidate molecules. Mol-E achieves state‑of‑the‑art performance on the Practical Molecular Optimization benchmark, scoring 17.500 in the task‑agnostic regime and 20.551 in the task‑informed regime. It also outperforms baseline methods in multi‑property optimization tasks involving docking against DRD2, MK2, and AChE.

By Philipp Guevorguian, Menua Bedrosian, Tigran Fahradyan, Gayane Chilingaryan, Armen Aghajanyan, Hrant Khachatrian
arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer