OpenAI Blog

Basis completes a tax workbook 2x faster with GPT-6 Astra

OpenAI Blog
5d ago

Introducing GPT-6.1 Sol

OpenAI introduces GPT‑6.1 Sol, a new model described as near‑Astra intelligence for coding, computer use, and professional work. It offers these capabilities at one‑fifth of Astra’s standard API input and output token prices, implying a more cost‑effective solution for developers and businesses.

arXiv AI
Sep 25

SheetMind: Actions Set Accuracy, Agents Set the Failure Mode

SheetMind is a Manager‑Action‑Reflection framework that evaluates how much spreadsheet agent performance derives from the agents themselves versus the shared action interface. In a controlled study on all 221 tasks of the SheetCopilot Benchmark, replacing the high‑level action API with primitive cell operations drops accuracy by 47.1 points, while adding a Reflection Agent improves performance by 4.5 points and a Manager by 1.4 points. The framework also shows that decomposition changes failure modes, reducing silent wrong outputs from 33% to 25%, and that GPT‑5 and GPT‑5‑mini achieve similar performance, whereas GPT‑3.5 underperforms significantly.

By Lyuhao Chen, Xi Cheng, Yanming Kang, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Yulang Fei, Brian Zhu, Daniel Jin, Binze Cai, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao
OpenAI Blog
Sep 22

Better prompt caching for GPT-6

The OpenAI Blog announces that GPT‑6 enhances prompt caching, achieving higher cache hit rates and introducing new diagnostics, explicit breakpoints, and controls. These features are designed to reduce latency and costs for users. The post highlights the technical improvements that make GPT‑6 more efficient and cost‑effective.

arXiv AI
Jul 3

Office Comprehension Benchmark

arXiv:2607. 01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.

By Firoz Shaik, Mateus Pican\c{c}o Lima Gomes, Tanvir Aumi, Jingci Wang, Milos Milunovic, Filip Basara, Ivana Jovanovic, Vishwas Suryanarayanan, Neha Nandan Kenkare, Weiyao Xie, Zhipeng Han, Zheng Zhang, Waleed Shahid, Jay Rathi, Russell Scherer, Thong Q. Nguyen, Michael Bentley, Tamara Stankovic, Rasika Chakravarthy, Vishal Chowdhary
OpenAI Blog
Sep 11

Cognition helps Devin test its own work with GPT‑6 Astra

The OpenAI Blog article titled "Cognition helps Devin test its own work with GPT‑6 Astra" discusses how GPT‑6 Astra enhances Devin’s capability to test software and demonstrate its functionality. This improvement aims to enable engineers to review less code and accelerate shipping of products.