arXiv Machine Learning By Yi Wu, Zheng Ren, Zhiyu Hu, Haochen Wang, Daryl Chang, Li Wei, Ting Wang, Zhen Li, Pooja Gupta, Nitin Jindal, Lukasz Heldt

Code-to-Harness: Distilling Black-Box Optimizers from Self-Play

Read the original on arXiv Machine Learning →

The paper investigates whether an agent can learn a numerical search strategy through executable practice and then encode that strategy as text. By repeatedly writing and evaluating optimizer programs, the agent distills a 197‑word text called Harness A, which significantly reduces regret for Gemini Flash and other language‑model executors, matching the performance of classical Gaussian‑process Bayesian optimization. An independent replication produced a different but equally effective text, Harness B, and the framework also achieved the lowest regret on a sealed YouTube reward‑tuning benchmark.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 26

Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search

arXiv:2606. 27291v1 Announce Type: new Abstract: Job-search platforms rely on low-bandwidth query interfaces that often fail to capture the high-dimensional complexity of candidate profiles.

By Ping Liu, Qianqi Shen, Jianqiang Shen, Wenqiong Liu, Rajat Arora, Yunxiang Ren, Chunnan Yao, Dan Xu, Baofen Zheng, Wanjun Jiang, Andrii Soviak, Kevin Kao, Jingwei Wu, Wenjing Zhang