arXiv Machine Learning

The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing

arXiv:2608. 04432v1 Announce Type: cross Abstract: On two-sided content platforms, symmetric two-sided isolation (assigning matched fractions of creators and viewers to isolated treatment and control submarkets) is widely used for creator-side and cold-start experiments because it removes cross-arm marketplace interference.

arXiv AI
Jun 9

Supracompetitive Pricing Under AI Monoculture

arXiv:2601. 01279v3 Announce Type: replace-cross Abstract: When competing sellers delegate pricing to a shared AI model, such as a large language model, correlated recommendations combined with performance-driven updates aggregating seller feedback raise a key question: can standard AI deployment practices inadvertently produce supracompetitive pricing?

By Shengyu Cao, Ming Hu
arXiv AI
Sep 10

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.

By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
arXiv AI
Sep 24

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.

By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv Machine Learning
Sep 11

The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems

The paper investigates whether an oligopolistic concentration of generative AI models accelerates or steers the phenomenon of model collapse when models are recursively trained on each other’s outputs. Using controlled ecosystems of 13 open‑source models and an injected probe that pushes one model’s market share to 90%, the authors find that varying market concentration has little effect on the speed or final state of collapse. Instead, the pace of collapse is largely determined by which models supply the training pool and how susceptible those models are to being carried along, with human‑written text in the pool roughly halving the drift.

By Yangze Liu, Zhongyi Han