Introducing AgentKit, new Evals, and RFT for agents
Today, we’re releasing new tools to help developers go from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents.
Related stories
Measuring Agents in Production
arXiv:2512. 04123v4 Announce Type: replace-cross Abstract: LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful.
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering.
The next evolution of the Agents SDK
OpenAI updates the Agents SDK with native sandbox execution and a model-native harness, helping developers build secure, long-running agents across files and tools.
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.
OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments
Reliable and Developer-Aligned Evaluation of Agents for Software Engineering
arXiv:2607. 06713v1 Announce Type: cross Abstract: Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments.
How can we assess human-agent interactions? Case studies in software agent design
arXiv:2510. 09801v3 Announce Type: replace Abstract: While benchmarks measure the accuracy of LLM-powered agents, they mostly assume full automation, failing to represent the collaborative nature of real-world use cases.
Harness engineering: leveraging Codex in an agent-first world
By Ryan Lopopolo, Member of the Technical Staff
ScreenSuite - The most comprehensive evaluation suite for GUI Agents!
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
Can Agents Design Libraries for Agents?
The paper introduces LibraryDesignBench, a benchmark that tests how well agents can design reusable libraries from specifications without prescribed designs. It evaluates libraries by measuring the correctness and simplicity of programs written by three different user agents across 242 programming problems in four languages. Findings show that while agents often replicate human-designed abstractions, downstream agents still tend to reimplement library features due to rigidity or usability issues, and that providing more prescriptive guidance improves reuse and program simplicity.