arXiv Machine Learning

Refine Drugs, Don't Complete Them: Uniform-Source Discrete Flows for Fragment-Based Drug Discovery

arXiv:2509. 26405v2 Announce Type: replace Abstract: We introduce InVirtuoGen, a discrete flow generative model for fragmented SMILES for de novo and fragment-constrained generation, and target-property/lead optimization of small molecules.

arXiv Machine Learning
Jun 10

Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors

arXiv:2602. 04119v2 Announce Type: replace Abstract: The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo that can be synthesized in practice.

By Hyeonah Kim, Minsu Kim, Celine Roget, Dionessa Biton, Louis Vaillancourt, Yves V. Brun, Yoshua Bengio, Alex Hernandez-Garcia
arXiv Machine Learning
Jul 21

Routing by Reaching: Composition of Pre-trained GFlowNets for Multi-Objective Generation

arXiv:2602. 21565v3 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) learn to sample diverse candidates in proportion to a reward function, making them well-suited for scientific discovery, where exploring multiple promising solutions is crucial.

By Seokwon Yoon, Youngbin Choi, Seunghyuk Cho, Seungbeom Lee, MoonJeong Park, Dongwoo Kim
arXiv Machine Learning
Jul 7

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.

By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
arXiv Machine Learning
Jun 25

Why Pool When You Can Flow? Active Learning with GFlowNets

arXiv:2509. 00704v2 Announce Type: replace Abstract: The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screening for drug discovery.

By Renfei Zhang, Mohit Pandey, Artem Cherkasov, Martin Ester