arXiv Machine Learning

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

arXiv Machine Learning
Jul 7

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.

By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
Hugging Face Trending Papers
Jul 29

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good.

arXiv Machine Learning
Sep 7

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

The paper explores how large language models (LLMs) can be trained for small-molecule drug design by using synthetic tasks that are cheaper to evaluate. By employing a curriculum that gradually increases task difficulty, the authors demonstrate that LLMs can learn design strategies that outperform larger models on structure-based lead optimization. This approach shows that scaling post‑training with synthetic tasks can effectively adapt LLMs to high‑cost experimental scenarios that are otherwise infeasible to train on directly.

By Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow
Hugging Face Trending Papers
Jun 17

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization

Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets. To overcome this supervision bottleneck, we introduce unsupervised reward optimization of PLMs, a comprehensive framework for steerable protein generation without ground-truth labels.