arXiv Machine Learning By Yue Pei, Hongming Zhang, Chao Gao, Martin M\"uller, Yingying Zhang, Mengxiao Zhu, Hao Sheng, Ziliang Chen, Liang Lin, Haogang Zhu

Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making

Read the original on arXiv Machine Learning →

arXiv:2508. 16420v3 Announce Type: replace Abstract: Target-conditioned sequence models provide a simple interface for controllable offline decision making, but the requested target return can be an unreliable control signal, especially when the target return lies in underrepresented regions of the dataset.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.