arXiv AI

HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning

HLS-Seek is a natural‑language‑to‑High‑Level‑Synthesis framework that optimizes Quality of Results (QoR) such as latency and resource usage by using a comparative proxy reward model instead of full synthesis‑in‑the‑loop reinforcement learning. The system achieves high syntax and functional correctness (84.7% and 81.4% respectively) with only 7 B parameters, surpasses GPT‑5.1 on functional pass@5, and trains 8.5× faster than real‑reward RL. In QoR evaluation, HLS‑Seek attains the lowest latency on 19 of 30 kernels and Pareto‑dominates baseline HLS tools on 9 kernels.

arXiv Machine Learning
Jun 26

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

arXiv:2606. 27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown.

By Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He