arXiv AI

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

arXiv:2607. 29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes.

arXiv Computation and Language
Aug 27

AEL: Evolving Agent Harness in Open-Ended Environments

The paper introduces Agent Evolving Learning (AEL), a two‑timescale framework that dynamically evolves an LLM agent’s memory‑retrieval harness in open‑ended environments. A fast Thompson‑Sampling bandit selects among retrieval policies each episode, while a slower LLM reflection diagnoses performance drops and injects new policies when the current set plateaus. AEL outperforms ten self‑improving and non‑LLM baselines on a sequential portfolio benchmark, boosting Sharpe ratio by 27% and achieving significant accuracy gains on a support‑ticket routing stream.

By Wujiang Xu, Jiaojiao Han, Minghao Guo, Kai Mei, Xi Zhu, Han Zhang, Dimitris N. Metaxas
arXiv AI
Sep 12

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA‑Skills is a new framework that treats skill optimization for large language model agents as a budgeted sequential problem over a dynamically evolving candidate set. It uses contextual‑bandit prioritization to focus evaluations on promising or informative candidates and refines the skill population based on execution feedback. In experiments across six agent benchmarks and three target models, COBRA‑Skills outperforms existing methods, cuts optimization cost by 55–58 % compared to SkillOpt, and requires only 50 unique optimization examples per benchmark.

By Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao, Zikun Qu, Junfeng Luo, Yao Shu, Bryan Kian Hsiang Low, Zhongxiang Dai
arXiv Computation and Language
Sep 3

CORAL: An LLM-Native Harness for Production Recommender Systems

CORAL is an LLM‑native harness that automates continual optimization of production recommender systems. It operates in a closed loop: an agent observes system signals, reasons over past decisions, and uses tools—including a numerical optimizer—to reconfigure the recommender while staying within a fixed operating budget. In A/B experiments on two large social platforms, CORAL improved engagement without extra serving cost on one platform and reduced serving cost without harming engagement on the other, demonstrating that a single agentic loop can replace manual engineering for ongoing system tuning.

By Muhammad Rafay Azhar, Yuhang Zhou, Gilbert Jiang, Yuchen Wang, Rahul Sharma, Matthew DeSousa, Jiayi Liu, Xin Guo, Lizhu Zhang, Xiangjun Fan
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv AI
Jul 24

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.

By Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
arXiv AI
Sep 25

Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits

CANOPY is a multi‑fidelity tree bandit algorithm that learns where a piecewise‑smooth prior holds instead of assuming global smoothness. It uses cheap random‑path probes to certify local aggregation bias and then focuses expensive leaf evaluations on cells where smoothness is violated. The method achieves provable fixed‑budget and regret guarantees that scale with the number of discontinuities, matching smooth‑tree rates when no violations exist and approaching structure‑blind search when violations are dense.

By Michael Jerge, Suman Jana