arXiv AI By Stefan Hut, Lorenzo Masoero

Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

Read the original on arXiv AI →

arXiv:2608. 02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

The paper introduces AgentX-Model, a dual‑agent framework that links proposal development with model experimentation in industrial recommender systems. The Research Agent drafts proposals from literature and prior findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework iteratively selects starting implementations and formulates new research questions, organizing work into Reproduce, Follow‑up, Composition, and Diagnose actions. Across production evaluations, most experiments exceeded business baselines, with recent A/B tests showing significant gains in acquisition efficiency, advertising spend, and watch time while reducing computational cost.

By Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou
arXiv Machine Learning
Sep 7

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

The paper introduces Speculative Uncertainty (SU), a technique that infers a failure likelihood for black‑box LLM agents by evaluating their generated token sequences with a lightweight draft model, without needing internal model details. SU extracts phase‑aware features from reasoning and action spans, calibrates them against verifiable outcomes, and produces a failure‑likelihood score usable by downstream policies. Applying a pre‑execution veto gate based on SU to software‑engineering agents such as Qwen3‑Coder‑480B and Claude 3.5 Sonnet reduced execution error rates by 6‑8 percentage points and token costs by 14‑19 %, while maintaining performance on out‑of‑distribution benchmarks and across different agent models.

By Konstantin Grotov, Valentin Malykh
Hugging Face Trending Papers
Sep 24

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems

Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems introduces AgentX-Model, a dual-agent framework that links proposal development with model experimentation in business-defined sandboxes. The Research Agent drafts proposals from literature and findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework cycles through Reproduce, Follow‑up, Composition, and Diagnose actions, achieving high AUC gains and significant business metric improvements in online A/B tests.

arXiv AI
Sep 25

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.

By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath