Agentic ML Exploration (A-MLE) for Ads Ranking
arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...
arXiv:2603. 22376v2 Announce Type: replace-cross Abstract: We present an AI Co-Scientist framework that closes the research loop for the production search-ranking system of a large online travel platform -- pairing LLM agents with direct cloud-compute access so that idea generation, code implementation, GPU experimentation, and result analysis iterate end-to-end with a human scientist in the loop.
arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...
arXiv:2606. 07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution.
Ideation Arena is a battle-style platform that evaluates research ideas generated by large language models (LLMs) and research agents through pairwise human assessment. The system builds shared literature contexts, collects over 6,000 double-blind comparisons from 105 computer science researchers, and constructs an Elo rating leaderboard to rank proposal-stage expert preferences. It also introduces Ideation Arena Eval, a benchmark to test whether automated evaluators align with human preferences, finding that current LLM judges achieve at best 72.56% Soft Accuracy on overall quality.
arXiv:2606. 31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.
arXiv:2606. 19704v1 Announce Type: new Abstract: Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes.
RankEvolve is an auto‑research framework that evolves generative ranking models by orchestrating multiple large‑language‑model coding agents through an Executable Operating Protocol (EOP). The system compiles a state machine that enforces phases, gates, branches, and loops, while a meta‑meta‑harness lets agents review and repair each other’s code. In budget‑matched experiments, heterogeneous composition of agents raised execution accuracy from 45.8 % to 62.5 % and reduced silent critical‑defect rates, achieving notable gains on the HSTU recommender and other benchmarks.
arXiv:2606. 07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify.
Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
arXiv:2607. 28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery.
The paper introduces Atria Dawn Preview, a foundation agentic language model aimed at scientific research and engineering workflows. Trained through a Verifiable Experience Pipeline, it performs competitively across 16 real‑world benchmarks, achieving the highest scores on five. The authors also present a detailed case study of human–AI collaboration, showing that while AI proposes methods and revisions, humans retain final decision‑making and guide the research direction.
arXiv:2607. 25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules.
arXiv:2608. 13940v1 Announce Type: new Abstract: AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days of GPU time.