Job-search platforms rely on low-bandwidth query interfaces that often fail to capture the high-dimensional complexity of candidate profiles. We present an end-to-end RLAIF (Reinforcement Learning from AI Feedback) framework to generate \emph{portable} job search queries, terms that abstract away seeker-specific identifiers while preserving generalizable qualifications.
arXiv:2607. 06854v1 Announce Type: cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade, since they beat a random opponent over 99 percent of the time and only tie copies of themselves.
By Nima Kelidari, Mohammadsaeed Haghi, Mahdi Salmani
arXiv:2604. 03501v5 Announce Type: replace-cross Abstract: Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those gains depend.
By Michael Caosun, Sinan Aral
arXiv:2606. 29457v1 Announce Type: new Abstract: When two companies bid to buy the same target, no one knows exactly what the target is worth.
By Zain Naboulsi
arXiv:2606. 27291v1 Announce Type: new Abstract: Job-search platforms rely on low-bandwidth query interfaces that often fail to capture the high-dimensional complexity of candidate profiles.
By Ping Liu, Qianqi Shen, Jianqiang Shen, Wenqiong Liu, Rajat Arora, Yunxiang Ren, Chunnan Yao, Dan Xu, Baofen Zheng, Wanjun Jiang, Andrii Soviak, Kevin Kao, Jingwei Wu, Wenjing Zhang
arXiv:2607. 06413v1 Announce Type: cross Abstract: Large language model coding agents increasingly perform open-ended data modeling and analysis.
By Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng
arXiv:2602. 12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove both individual and group outcomes.
By Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian
The paper investigates the limitations of post-training AI agents that can autonomously train large language models. It distinguishes between execution-level capability—making adjustments within a chosen training strategy—and strategy-level capability—revising the overall approach based on new evidence. Analysis of many public post-training runs shows that agents lock into a strategy early and then only perform local tweaks, regardless of task. Experiments with experience scaffolds, human guidance, and extra compute improve execution but do not enable strategy reevaluation, indicating that agents lack a mechanism to spontaneously reassess their strategy during training.
By Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin
arXiv:2512. 04988v2 Announce Type: replace-cross Abstract: Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agentic swarms.
By Christopher Chiu, Simpson Zhang, Mihaela van der Schaar
arXiv:2608.30047v1 Announce Type: new
Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they...
By Shitanshu Bhushan, Yunxiang Zhang, Lu Wang
arXiv:2608.30724v1 Announce Type: cross
Abstract: LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented...
By Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs, Julian Moncarz, Kaustubh Kislay, Juan J. Vazquez
arXiv:2609.36308v1 Announce Type: new
Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In rece...
By Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks