arXiv Machine Learning By Xinyu Wang, Jinbo Bi, Minghu Song

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Read the original on arXiv Machine Learning →

arXiv:2607. 26391v1 Announce Type: new Abstract: Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 29

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good.

arXiv Machine Learning
Jul 7

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.

By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue