Introducing GPT-5.1 for developers
GPT-5. 1 is now available in the API, bringing faster adaptive reasoning, extended prompt caching, improved coding performance, and new apply_patch and shell tools.
The OpenAI Blog announces that GPT‑6 enhances prompt caching, achieving higher cache hit rates and introducing new diagnostics, explicit breakpoints, and controls. These features are designed to reduce latency and costs for users. The post highlights the technical improvements that make GPT‑6 more efficient and cost‑effective.
GPT-5. 1 is now available in the API, bringing faster adaptive reasoning, extended prompt caching, improved coding performance, and new apply_patch and shell tools.
arXiv:2607. 15516v1 Announce Type: cross Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent).
arXiv:2609.16215v1 Announce Type: new Abstract: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate s...
ChatGPT introduces improved GPT-5. 6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.
See how a group of leading developers use GPT-5 for the first time.
The article explains how startups can select GPT‑6 models and adjust their reasoning effort. It offers guidance on improving prompts and skills, coordinating tools, and preparing production workflows.
The paper introduces a cache‑aware post‑training framework for Mixture‑of‑Experts (MoE) models that jointly adapts the MoE backbone and lightweight auxiliary cache routers while keeping the native Top‑K expert‑selection rule. Two modes are proposed: Temporal Router, which predicts same‑layer reuse and retains experts for future tokens, and Spatio‑Temporal Router, which adds a Spatio Router that refines the temporal cache using the causal predecessor’s hidden state. Experiments on Qwen3 and GPT‑OSS across GSM8K, MATH, and CommonsenseQA show that Temporal Router improves cache hit rates and reduces expert‑weight traffic, while Spatio‑Temporal Router achieves the best load‑adjusted efficiency, outperforming strong prefetching baselines.
PrefixBench-H100 is a reproducible benchmark that evaluates how reusing prompt prefixes affects LLM serving performance on NVIDIA H100 GPUs. It tests two popular runtimes (vLLM and TensorRT-LLM) across varied workloads, measuring metrics such as time-to-first-token, latency, throughput, cache hits, and GPU memory usage. The study identifies when prefix reuse significantly reduces first‑token latency and when cache pressure diminishes those gains, noting that cache effectiveness is largely unaffected by concurrency or output length, while differences arise mainly in scheduling.
Introducing GPT-5 in our API platform—offering high reasoning performance, new controls for devs, and best-in-class results on real coding tasks.
The paper evaluates a learned request‑routing policy for disaggregated large‑language‑model serving, where compute‑heavy prefill and memory‑heavy decode stages run on separate GPU pools. Using a discrete‑event simulator and real NVIDIA A40 GPUs, the calibrated router—leveraging prompt length, predicted output length, KV‑cache pressure, and SLO class—outperforms round‑robin, least‑loaded, and length‑based heuristics, achieving the highest mean goodput (0.864) and lowest variance across three mixed, bursty arrival traces. Hardware calibration proves critical, providing a 4.5‑point goodput boost and roughly 40 % of the tail‑latency advantage, and the learned router can match round‑robin performance with one fewer GPU in certain scenarios.