OpenAI Blog

New tools for building agents

Towards Data Science
Aug 4

Using Agents as Tools

Building manager–specialist workflows with the OpenAI Agents SDK The post Using Agents as Tools appeared first on Towards Data Science .

By Shuai Guo
arXiv AI
Jun 8

Measuring Agents in Production

arXiv:2512. 04123v4 Announce Type: replace-cross Abstract: LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful.

By Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis
arXiv AI
Sep 11

A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model

The paper presents a taxonomy of architecture options for foundation-model-based agents, covering functional capabilities, non‑functional qualities, and operational aspects of design‑time and run‑time phases. It also introduces a decision model to guide critical design and runtime choices, aiming to streamline and improve the development of such agents. By unifying these classifications, the authors seek to reduce fragmentation in the field and provide a structured framework for architects and developers.

By Jingwen Zhou, Qinghua Lu, Jieshan Chen, Liming Zhu, Xiwei Xu, Zhenchang Xing, Stefan Harrer
arXiv AI
Jul 7

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.

By Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
4d ago

Can Agents Design Libraries for Agents?

The paper introduces LibraryDesignBench, a benchmark that tests how well agents can design reusable libraries from specifications without prescribed designs. It evaluates libraries by measuring the correctness and simplicity of programs written by three different user agents across 242 programming problems in four languages. Findings show that while agents often replicate human-designed abstractions, downstream agents still tend to reimplement library features due to rigidity or usability issues, and that providing more prescriptive guidance improves reuse and program simplicity.

By Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt