Introducing the SWE-Lancer benchmark
Read the original on OpenAI Blog →Can frontier LLMs earn $1 million from real-world freelance software engineering?
Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.
Can frontier LLMs earn $1 million from real-world freelance software engineering?
Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.
arXiv:2607. 14108v1 Announce Type: cross Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory.
arXiv:2601. 06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications.
arXiv:2606. 27406v1 Announce Type: cross Abstract: Software engineering, whether performed by humans or by AI agents, requires reasoning about how software behaves.
OpenAI’s B2B Signals research shows how frontier enterprises deepen AI adoption, scale Codex-powered agentic workflows, and build durable competitive advantage.
arXiv:2606. 05548v1 Announce Type: cross Abstract: The rapid proliferation of Agent Development Kits (ADKs), SDK-level frameworks for building LLM-powered autonomous agents, has outpaced any empirical understanding of how framework choice affects agent performance.
SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage.