E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.
arXiv:2608.29843v1 Announce Type: cross
Abstract: Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurem...
By Louis Yiven Zhu
arXiv:2606. 18543v1 Announce Type: new Abstract: Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.
By Haozhe Chen, Karthik Narasimhan, Zhuang Liu
Analysis provides insights into ChatGPT’s impact on the economy. OpenAI also launches new research collaboration to study AI’s broader effects on the labor market and productivity.
arXiv:2607. 20349v1 Announce Type: cross Abstract: Generative AI can produce book-length works of fiction at near-zero cost.
By Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon
arXiv:2607. 07207v1 Announce Type: cross Abstract: We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.
By Satoshi Matsuoka
OpenAI’s business model scales with intelligence—spanning subscriptions, API, ads, commerce, and compute—driven by deepening ChatGPT adoption.
arXiv:2509. 26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions.
By Berdymyrat Ovezmyradov
LY Corporation: Driving growth and ‘WOW’ moments with OpenAI
The article discusses Anthropic’s Claude Fable 5.1 release, highlighting its claimed improvements in coding, knowledge work, and problem‑solving, particularly a 52.6% score on the new Terminal‑Bench‑Science 0.1 benchmark. The author examines the model’s performance on the pelican benchmark, noting that Fable 5.1’s five reasoning levels (low, medium, high, xhigh, max) sometimes skip reasoning entirely for certain prompts, as evidenced by token counts and cost metrics. The piece provides detailed transcript data for each reasoning level when generating an SVG of a pelican riding a bicycle.
arXiv:2608. 08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months.
By Jan Sp\"orer