arXiv:2511. 03836v2 Announce Type: replace Abstract: Deep Q-Networks (DQNs) estimate future returns by learning from transitions sampled from a replay buffer.
By Lipeng Zu, Hansong Zhou, Xiaonan Zhang
arXiv:2609.08232v1 Announce Type: cross
Abstract: Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of design rules. Modern routers can struggle t...
By Afsara Khan, Austin Rovinski
arXiv:2408.09838v3 Announce Type: replace
Abstract: A continual learning agent builds on previous experiences to develop increasingly complex behaviors by adapting to non-stationary and dynamic envir...
By Achref Jaziri, Etienne K\"unzel, Visvanathan Ramesh
The paper introduces a history‑aware offline reinforcement learning policy that predicts iterative cost weights for routing in dense integrated circuit designs. By incorporating a lightweight LSTM and additional router features, the policy retains sequence context and improves convergence across various placement densities and guide qualities. Integrated into any cost‑based router with minimal changes, the approach reduces design rule violations by an average of 92% and cuts runtime by 10%.
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function.
arXiv:2606. 00350v1 Announce Type: cross Abstract: Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates.
By Anas Houssaini, Mohamad H. Danesh, Amin Abyaneh, Scott Fujimoto, Hsiu-Chin Lin, David Meger