arXiv AI

Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

arXiv:2608. 02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry.

arXiv AI
Sep 25

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

The paper introduces AgentX-Model, a dual‑agent framework that links proposal development with model experimentation in industrial recommender systems. The Research Agent drafts proposals from literature and prior findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework iteratively selects starting implementations and formulates new research questions, organizing work into Reproduce, Follow‑up, Composition, and Diagnose actions. Across production evaluations, most experiments exceeded business baselines, with recent A/B tests showing significant gains in acquisition efficiency, advertising spend, and watch time while reducing computational cost.

By Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou
arXiv Machine Learning
Sep 7

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

The paper introduces Speculative Uncertainty (SU), a technique that infers a failure likelihood for black‑box LLM agents by evaluating their generated token sequences with a lightweight draft model, without needing internal model details. SU extracts phase‑aware features from reasoning and action spans, calibrates them against verifiable outcomes, and produces a failure‑likelihood score usable by downstream policies. Applying a pre‑execution veto gate based on SU to software‑engineering agents such as Qwen3‑Coder‑480B and Claude 3.5 Sonnet reduced execution error rates by 6‑8 percentage points and token costs by 14‑19 %, while maintaining performance on out‑of‑distribution benchmarks and across different agent models.

By Konstantin Grotov, Valentin Malykh
Hugging Face Trending Papers
Sep 24

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems introduces AgentX-Model, a dual-agent framework that links proposal development with model experimentation in business-defined sandboxes. The Research Agent drafts proposals from literature and findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework cycles through Reproduce, Follow‑up, Composition, and Diagnose actions, achieving high AUC gains and significant business metric improvements in online A/B tests.

arXiv AI
Sep 25

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.

By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv Machine Learning
Sep 23

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.

By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng
arXiv AI
Sep 25

When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.

By Yufeng Wang
arXiv Machine Learning
5d ago

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

The paper explores how data from fixed A/B tests can guide the deployment of adaptive experiments using contextual bandits. By combining off‑policy evaluation with a controlled warm‑start simulation, the authors rank pre‑specified adaptive and non‑adaptive policies using doubly robust estimators. Experiments on synthetic trials and real benchmarks show that adaptive, context‑aware policies outperform fixed allocations when heterogeneity exists, but offer little advantage otherwise.

By Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik