arXiv Machine Learning By Yuche Gao, Jos\'e Miguel Hern\'andez-Lobato, Siyuan Guo

PerturbPFN: Probing the Limits of Synthetic Priors in Drug Perturbation Modelling

Read the original on arXiv Machine Learning →

arXiv:2607. 23447v1 Announce Type: new Abstract: Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv Machine Learning
1d ago

A Large Scale Investigation of Scaling Limits in Chemical Language Models

The paper reports a large-scale, compute-controlled study of Chemical Language Models (CLMs) involving over 30,000 experiments across different molecular representations, tokenizations, model sizes, datasets, and architectures. It finds clear scaling trends in pretraining loss but shows that these improvements do not translate into proportional gains in goal-directed molecular design, with chemical syntax saturating early while semantic properties develop more slowly. The authors release a new suite of models, NovoMolGen, that achieves state-of-the-art results in drug discovery tasks, highlighting a disconnect between representation learning and downstream design and calling for new pretraining paradigms that target chemical semantics.

By Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar
arXiv Machine Learning
Aug 24

PerturbRx: Learning Treatment-Conditioned Latent Transitions for Patient Drug Response Prediction

PerturbRx is a treatment‑conditioned representation learning framework that learns latent transitions induced by drug interventions. It trains a drug‑ and dose‑conditioned transition predictor using control and treated single‑cell populations, then applies this predictor to pretreatment patient profiles to generate response features without needing post‑treatment data. On TCGA and patient‑derived xenograft benchmarks, PerturbRx outperforms other methods, demonstrating the value of perturbation‑pretrained latent transitions for patient‑level drug‑response prediction.

By Yoshitaka Inoue, Minoh Jeong, Alfred Hero, Rui Kuang, Augustin Luna
arXiv AI
Aug 25

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA is a scalable multimodal framework that learns molecular world models by using modality masking instead of suboptimal perturbations. It incorporates diverse data such as molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations, and other drug‑discovery information. Benchmarks show that the representations it learns perform strongly, highlighting the benefit of embedding biochemical context via latent‑space prediction.

By Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff