Function calling and other API updates
We’re announcing updates including more steerable API models, function calling capabilities, longer context, and lower prices.
Increasing enterprise support with more security features and controls, updates to our Assistants API, and tools to better manage costs.
We’re announcing updates including more steerable API models, function calling capabilities, longer context, and lower prices.
arXiv:2607. 11348v1 Announce Type: cross Abstract: AI code assistants are transforming software development, but their implications for software security remain a major concern, particularly in the context of security APIs.
API Partnership with Stack Overflow Stack Overflow and OpenAI today announced a new API partnership that will empower developers with the collective strengths of the world’s leading knowledge platform for highly technical content with the world’s most popular LLM models for AI development.
Decagon and OpenAI deliver high-performance, fully automated customer support at scale
arXiv:2607. 25042v1 Announce Type: new Abstract: The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when accessing enterprise data without predefined API endpoints.
We’re adding new features to help developers have more control over fine-tuning and announcing new ways to build custom models with OpenAI.
arXiv:2605. 10555v2 Announce Type: replace Abstract: As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-oriented CRUD paradigms.
Introducing OpenAI o1, Realtime API improvements, a new fine-tuning method and more for developers.
arXiv:2604. 00073v3 Announce Type: replace-cross Abstract: There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomously.
Salesforce integrates OpenAI’s enterprise-ready LLMs to transform customer applications.
The paper examines whether benchmark scores obtained via APIs accurately reflect the performance of AI chatbots when accessed through user interfaces. By auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks, the authors find that API evaluations consistently overestimate accuracy and consistency compared to interface evaluations, with differences comparable to downgrading a full model generation. Attempts to align API behavior with interface behavior through prompt and parameter adjustments only partially close the gap.