E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.
arXiv:2608.29843v1 Announce Type: cross
Abstract: Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurem...
By Louis Yiven Zhu
arXiv:2606. 18543v1 Announce Type: new Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.
By Haozhe Chen, Karthik Narasimhan, Zhuang Liu
Analysis provides insights into ChatGPT’s impact on the economy. OpenAI also launches new research collaboration to study AI’s broader effects on the labor market and productivity.
arXiv:2607. 20349v1 Announce Type: cross Abstract: Generative AI can produce book-length works of fiction at near-zero cost.
By Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon