arXiv Machine Learning

RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning

arXiv AI
Aug 19

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

The paper introduces MCTH (Monte Carlo Tree Hallucination), an inference-only framework that performs all‑atom biomolecular sequence‑structure co‑design by treating pretrained folding and inverse‑folding models as black‑box operators. MCTH uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence, uncertainty, and cross‑expert consensus. Experiments across protein‑RNA, protein‑DNA, protein‑protein, and protein‑ligand design show that adaptive search outperforms simpler sampling strategies, and evaluations with AlphaFold3 and Chai‑1 demonstrate transferability beyond the search‑time oracle.

By Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens
arXiv Machine Learning
Aug 21

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

arXiv:2506. 07459v4 Announce Type: replace Abstract: Protein generative models have shown remarkable promise in protein design, yet their success rates remain constrained by reliance on curated sequence-structure datasets and by misalignment between supervised objectives and real design goals.

By Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu
Hugging Face Trending Papers
Aug 19

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties such as drug‑likeness or binding affinity while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. The method addresses a reward‑hacking failure mode by treating each input molecule as the start of a forward‑synthesis trajectory, applying learned reaction templates with in‑stock building blocks, and producing diverse, high‑quality outputs with explicit synthesis routes.