arXiv Machine Learning By Kseniia Vaniushkina, Jeongmin Lim, Jinyong Park

GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures

Read the original on arXiv Machine Learning →

arXiv:2607. 17671v1 Announce Type: new Abstract: Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
Jul 29

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.

By Kai Lun Huang (California State University, Fullerton), Wei Chieh Sun (University of Washington)