arXiv Machine Learning By Steve Hanneke, Hongao Wang, Mingyue Xu

Towards a theory of inference-time alignment with unknown rewards

Read the original on arXiv Machine Learning →

arXiv:2608. 15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.