arXiv Machine Learning By Tshemollo Rapolai, Seite Makgai, Mohammad Arashi

Feature Space Selection and Heterogeneous Effect Estimation for Blood-Brain Barrier Permeability: A Random Forest to the Generalized Random Forest Pipeline

Read the original on arXiv Machine Learning →

The study evaluates how different molecular feature spaces—Morgan fingerprints, RDKit physicochemical descriptors, and SMILES bigrams—affect the prediction of blood‑brain barrier permeability using various learning algorithms. Dynamic Random Forests with combined features achieved the best performance (mean AUC 0.970). When applying Generalized Random Forests to estimate heterogeneous effects of LogP on BBB permeability, orthogonalization revealed that apparent heterogeneity largely vanished after accounting for confounding, shifting importance toward residual SMILES bigram information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 6

Geometry-Informed Parameter-Efficient Fine-Tuning of Pre-trained Molecular GNNs for Blood-Brain Barrier Permeability Prediction

arXiv:2608. 04257v1 Announce Type: new Abstract: Blood-brain barrier permeability (BBBP) prediction is a critical screening task in central nervous system drug discovery, where candidate molecules must be assessed for whether they can cross, or should be prevented from crossing, the blood-brain barrier.

By Marco Vieto Vega, Long D. Nguyen, Binh P. Nguyen
arXiv Machine Learning
Sep 16

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

The paper introduces a new algorithm that uses decision trees and random forests to estimate individual treatment effects while providing interpretability. It modifies the standard random forest splitting criterion by combining a heterogeneity-focused criterion with a bias-correction criterion, enabling the model to handle observational studies with varying treatment propensities without separately estimating propensity scores. The resulting tree structure directly reveals which features drive treatment effect differences, and simulation studies show the method matches or surpasses existing approaches in prediction accuracy while improving interpretability.

By Nicolas Alexander Ihlo, Merle Behr
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon