arXiv:2607. 10202v1 Announce Type: new Abstract: Cross-model comparisons read divergence in value dispositions as evidence that language models hold individuated values.
By Hong-In Won, Jinseok Jang, Hyoseop Kim
arXiv:2608.29430v1 Announce Type: cross
Abstract: Industrial recommenders give new content initial views through budgeted exploration, then use early performance to decide further delivery. On many s...
By Yuanyuan Shen, Yiren Yan, Wenjie Li, Chunhui Zhu
arXiv:2601. 01279v3 Announce Type: replace-cross Abstract: When competing sellers delegate pricing to a shared AI model, such as a large language model, correlated recommendations combined with performance-driven updates aggregating seller feedback raise a key question: can standard AI deployment practices inadvertently produce supracompetitive pricing?
By Shengyu Cao, Ming Hu
arXiv:2607. 25655v1 Announce Type: new Abstract: Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.
By Jesung Park
The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.
By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
arXiv:2602. 16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments.
By Zehao Xu, Tony Paek, Kevin O'Sullivan, Attila Dobi
arXiv:2608. 08395v1 Announce Type: cross Abstract: Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents.
By Lingxiu Dong, Kaiwen Luo, Fasheng Xu
CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.
By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
The paper investigates whether an oligopolistic concentration of generative AI models accelerates or steers the phenomenon of model collapse when models are recursively trained on each other’s outputs. Using controlled ecosystems of 13 open‑source models and an injected probe that pushes one model’s market share to 90%, the authors find that varying market concentration has little effect on the speed or final state of collapse. Instead, the pace of collapse is largely determined by which models supply the training pool and how susceptible those models are to being carried along, with human‑written text in the pool roughly halving the drift.
By Yangze Liu, Zhongyi Han
Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16. 1M occurrences), human results are not balanced.
arXiv:2608. 12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it.
By Binshuang Li
arXiv:2607. 18045v1 Announce Type: new Abstract: Organizations often pool dispersed information into one ranking and then allow many agents to act on that shared view.
By Yohei Nakajima