arXiv:2606. 03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward.
By Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Zaheer Abbas, Eser Ayg\"un, David Smalling, Shibl Mourad, Doina Precup, Andr\'e Barreto, Mark Rowland
arXiv:2606. 24990v1 Announce Type: new Abstract: Reinforcement Learning (RL) has become a powerful paradigm for de novo molecular design, enabling Chemical Language Models (CLMs) to navigate and explore the chemical space while optimizing specific desired properties.
By Borja Medina, Jon Paul Janet
The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.
By Oussama Boussif, L\'ena N\'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity.
arXiv:2606. 26657v1 Announce Type: new Abstract: Identifying high-utility candidates from massive discrete spaces under expensive evaluations is a recurring challenge across the sciences, with structure-based drug discovery as a prominent example.
By Mohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay, Joana Reis, Felix Strieth-Kalthoff, Haribabu Arthanari
PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. Experiments demonstrate that PGFS++ enhances target properties and preserves high output diversity, overcoming the reward‑hacking failure mode seen in earlier versions.
By Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
arXiv:2607. 19044v1 Announce Type: new Abstract: Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design.
By Mingxuan Ouyang, Hao Lan, Wanyu Lin
arXiv:2607. 00531v1 Announce Type: cross Abstract: Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge.
By Xuefeng Liu, Mingxuan Cao, Qinan Huang, Thomas Brettin, Rick Stevens, Le Cong
PGFS++ is a synthesis‑aware reinforcement learning framework that improves molecular properties such as drug‑likeness or binding affinity while ensuring the resulting molecules can be synthesized and remain structurally similar to the input. It builds on PGFS+ by using trainable embedding lookup tables for reaction templates and second reactants, a more effective scoring function, and a refined RL algorithm. The method addresses a reward‑hacking failure mode by treating each input molecule as the start of a forward‑synthesis trajectory, applying learned reaction templates with in‑stock building blocks, and producing diverse, high‑quality outputs with explicit synthesis routes.
arXiv:2507. 15356v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions.
By Lu Guo, Yixiang Shan, Zhengbang Zhu, Qifan Liang, Lichang Song, Ting Long, Weinan Zhang, Yi Chang
arXiv:2605. 01248v3 Announce Type: replace Abstract: Reinforcement learning (RL) post-training has enabled newer capabilities in models, such as agentic tool-use for search.
By Harsh Goel, Akhil Udathu, Susmija Jabbireddy, Pradnesh Kalkar, Atharva Parulekar
arXiv:2505. 15201v5 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) algorithms sample multiple n>1 solution attempts for each problem and reward them independently.
By Christian Walder, Deep Karkhanis