arXiv:2602. 04119v2 Announce Type: replace Abstract: The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo that can be synthesized in practice.
By Hyeonah Kim, Minsu Kim, Celine Roget, Dionessa Biton, Louis Vaillancourt, Yves V. Brun, Yoshua Bengio, Alex Hernandez-Garcia
arXiv:2606. 07610v1 Announce Type: cross Abstract: State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response.
By Argyrios Gerogiannis, Yekaterina Yegorova, Mark Hasegawa-Johnson, Venugopal V. Veeravalli
The paper introduces a new learning objective called trajectory balance for Generative Flow Networks (GFlowNets), aiming to improve credit assignment across long action sequences. It demonstrates that minimizing this objective yields a policy that samples exactly from the target distribution. Experiments on four domains show that trajectory balance enhances convergence, sample diversity, and robustness to long sequences and large action spaces.
By Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio, Chen Sun, Yoshua Bengio
Selective Regenerative Decoding (SRD) is a new inference-time decoding method that improves large language model reasoning by allowing segment-level intervention on candidate trajectories. Instead of discarding or keeping entire trajectories, SRD selectively refines only the degraded suffix while preserving useful prefixes, leading to higher expected trajectory quality and better sample efficiency. Experiments on MATH500, GPQA Diamond, HotpotQA, and AlpacaEval show that SRD matches Best-of-N accuracy with fewer generated tokens and outperforms speculative rejection in low‑compute settings.
By Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung, Yi-an Lai, Arshit Gupta
The paper introduces the Implicit Prefix-Value Reward Model (IPVRM), which learns the probability of eventual correctness for each prefix directly from outcome labels, thereby aligning training targets with inference-time step signals via temporal-difference differences. IPVRM improves step-verification F1 on ProcessBench. Additionally, the authors propose Distribution-Level RL (DistRL), a policy optimization method that applies TD advantages to both sampled and high-probability tokens, offering dense counterfactual updates without extra rollouts, and show that DistRL consistently enhances downstream reasoning when combined with IPVRM.
By Shiping Gao, Hongzhan Chen, Xiaojun Quan, Qifan Wang, Lifu Huang
arXiv:2602. 21565v3 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) learn to sample diverse candidates in proportion to a reward function, making them well-suited for scientific discovery, where exploring multiple promising solutions is crucial.
By Seokwon Yoon, Youngbin Choi, Seunghyuk Cho, Seungbeom Lee, MoonJeong Park, Dongwoo Kim