arXiv Machine Learning

GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures

arXiv:2607. 17671v1 Announce Type: new Abstract: Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response?

arXiv AI
Sep 25

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of existing datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys. The benchmark demonstrates that performance drops sharply when moving from random‑decoy to hard‑negative evaluation, and it releases data, splits, code, and baseline implementations for reproducible comparison.

By Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang, Baris Coskunuzer
Hugging Face Trending Papers
Sep 24

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

TopU-LBVS is a new multi‑target benchmark for ligand‑based virtual screening that addresses shortcomings of previous datasets by using hard‑negative decoys and a fixed 1:40 active‑to‑decoy ratio. It covers 93 protein targets across seven classes, provides three evaluation protocols (full, low‑data, and mini), and includes curated ChEMBL‑35 bioactivity data with property‑matched, structurally similar decoys to reduce shortcut learning. The benchmark comes with released data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of LBVS and molecular representation methods.

arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
Jul 29

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

arXiv:2607. 24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution.

By Kai Lun Huang (California State University, Fullerton), Wei Chieh Sun (University of Washington)
Hugging Face Trending Papers
Sep 2

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.

arXiv AI
Sep 3

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI is a new framework for drug‑target interaction prediction that uses probe‑driven pattern matching to preserve weak biochemical signals. It introduces IterProbe, which retains contextual states across refinement depths and selects them with learnable probes, and BindingProbe, which models drug‑protein complementarity at both local and whole‑pair levels. Experiments show that ProbeMatchDTI outperforms existing methods, improving AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank, and its predictions can be integrated into downstream drug‑discovery workflows.

By Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
Hugging Face Trending Papers
Jul 6

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles. Preclinical transfer-learning models can simulate drug-induced expression changes but are often hard to interpret and unstable, whereas knowledge-graph methods provide mechanistic context yet remain static and fail to capture drug-induced transcriptomic perturbation dynamics.

arXiv Machine Learning
Sep 22

DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction

The paper introduces DPTM‑DT, a dual‑pretrained Transformer framework that integrates GROVER molecular graph embeddings, ESM protein language‑model embeddings, and CTD physicochemical descriptors for drug‑target prediction. It employs bidirectional cross‑modal attention to share drug‑target information and uses a single pair representation for continuous affinity regression, high‑affinity binary classification, and six‑level affinity classification. Experiments on Davis and KIBA datasets show that DPTM‑DT outperforms existing methods across regression, binary, and multiclass tasks, with ablation studies confirming the contributions of dual target representation, gated fusion, and cross‑modal attention.

By Ge Kong
arXiv Machine Learning
Sep 24

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.

By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang