arXiv Machine Learning

Feature Space Selection and Heterogeneous Effect Estimation for Blood-Brain Barrier Permeability: A Random Forest to the Generalized Random Forest Pipeline

The study evaluates how different molecular feature spaces—Morgan fingerprints, RDKit physicochemical descriptors, and SMILES bigrams—affect the prediction of blood‑brain barrier permeability using various learning algorithms. Dynamic Random Forests with combined features achieved the best performance (mean AUC 0.970). When applying Generalized Random Forests to estimate heterogeneous effects of LogP on BBB permeability, orthogonalization revealed that apparent heterogeneity largely vanished after accounting for confounding, shifting importance toward residual SMILES bigram information.

arXiv Machine Learning
Aug 6

Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs for Blood-Brain Barrier Permeability Prediction

arXiv:2608. 04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing, the blood-brain barrier.

By Marco Vieto Vega, Long D. Nguyen, Binh P. Nguyen
arXiv Machine Learning
Sep 16

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.

By Nicolas Alexander Ihlo, Merle Behr
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
Hugging Face Trending Papers
Sep 2

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI introduces a probe-driven framework for drug‑target interaction prediction that preserves weak biochemical signals by using IterProbe to retain contextual states and BindingProbe to model cross‑entity complementarity at multiple scales. The method improves AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank compared to prior biochemical representation learning approaches. Feature‑level analyses confirm the effectiveness of the probe-driven pattern matching, and the predictions are linked to an evidence‑guided downstream drug‑discovery workflow for candidate refinement and validation planning.

arXiv AI
Sep 3

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI is a new framework for drug‑target interaction prediction that uses probe‑driven pattern matching to preserve weak biochemical signals. It introduces IterProbe, which retains contextual states across refinement depths and selects them with learnable probes, and BindingProbe, which models drug‑protein complementarity at both local and whole‑pair levels. Experiments show that ProbeMatchDTI outperforms existing methods, improving AUC‑ROC by 2.0% on BindingDB and 0.5% on DrugBank, and its predictions can be integrated into downstream drug‑discovery workflows.

By Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
arXiv Machine Learning
Aug 31

Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining

The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.

By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz
arXiv Machine Learning
Jun 9

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

arXiv:2604. 26498v3 Announce Type: replace Abstract: The rapid growth of molecular foundation models and large language models (LLMs) has encouraged a scale centred view of AI in drug discovery, in which larger pretrained models are expected to supersede compact cheminformatics models.

By Jinjiang Guo, Sheng Ding
arXiv Machine Learning
3d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola