arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
By Davy Guan, Lu Zhang, Asiri Wijesinghe, Allen Zhu, He Zhao, Helen Power, F. Hafna Ahmed, Andrew Warden, Cheng Soon Ong, Daniel M. Steinberg
arXiv:2609.38744v1 Announce Type: new
Abstract: Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials de...
By Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo
The paper introduces a new Bayesian optimization approach tailored for generative models used in de novo discovery pipelines. By employing a linear surrogate model constrained to a spherical domain—where high‑dimensional latent vectors naturally concentrate—the authors derive nearly closed‑form solutions for both surrogate modeling and acquisition, achieving at least a 100‑fold speedup over existing methods. This acceleration enables Bayesian optimization to be used as a practical drop‑in component in pipelines that previously found it too slow to consider.
By Donney Fan, Colin Doumont, Aleksandra Kalisz, Paul Duckworth, Jacob R. Gardner, Henry Moss, Geoff Pleiss
arXiv:2608.22967v1 Announce Type: new
Abstract: Practical molecular inverse design is rarely a one-shot generation problem; it often takes the form of closed-loop candidate-pool enrichment, where und...
By Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Lei Bai, Tianshu Yu
Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.
By Blazej Banaszewski, Andrew W. Fitzgibbon
arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.
By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
The paper investigates nonlinear dimensionality reduction for Bayesian optimisation (BO) by transforming high‑dimensional black‑box optimisation problems into a sequence of low‑dimensional latent‑space BO (LSBO) tasks. It extends earlier linear embedding approaches by using variational autoencoders (VAEs), deep metric loss, and adaptive retraining to better capture nonlinear structure, and couples LSBO with sequential domain reduction (SDR‑LSBO) to progressively narrow search domains. Experiments on GPU‑accelerated BoTorch with Matérn‑5/2 Gaussian‑process surrogates show that VAE‑based LSBO outperforms adaptive linear embeddings, and the authors provide a theoretical analysis of latent‑space error versus representation gap under PAC‑Bayes conditions.
By Luo Long, Coralia Cartis, Paz Fink Shustin
arXiv:2606. 30258v1 Announce Type: cross Abstract: Tabular foundation models have advanced deep learning for tabular data by delivering strong default performance across many small and medium tasks.
By Boshko Koloski, Xiangjian Jiang, Senja Pollak, Bla\v{z} \v{S}krlj, Mateja Jamnik, Nikola Simidjievski
arXiv:2608. 04113v1 Announce Type: cross Abstract: Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lower-fidelity proxies available.
By Gustavo Sutter, Hao Wang, Luis Ricardez-Sandoval, Pascal Poupart, Agustinus Kristiadi
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis
arXiv:2607. 23447v1 Announce Type: new Abstract: Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space.
By Yuche Gao, Jos\'e Miguel Hern\'andez-Lobato, Siyuan Guo
arXiv:2607. 23480v1 Announce Type: new Abstract: Variational autoencoders (VAEs) transform high-dimensional, often noisy data into a compact latent representation, making downstream optimization more tractable.
By Ye Shi