Writing
September 24, 2026 · 6 min read

OpenAI tightens prompt caching for GPT-6, and it matters more than it sounds

OpenAI's GPT-6 prompt caching update raises hit rates, adds explicit cache breakpoints and a diagnostics dashboard, and pushes cached-token discounts to as much as 90% — a plumbing change that quietly reshapes how you should design long-running agents.

openaigpt-6prompt-cachingllm-infrastructureagentic-ai

OpenAI shipped an update to prompt caching for GPT-6: higher default cache hit rates, explicit cache breakpoints, a dashboard for cache performance, a diagnostics tool that names why a request missed cache, and discounts of up to 90% on cached input tokens for shared prefixes reused within a 30-minute window. No new model, no new capability. Just the plumbing underneath every API call getting less wasteful. I want to explain why that plumbing is worth your attention.

What actually changed

Prompt caching isn't new — OpenAI has discounted repeated prefixes since 2024. What's new in this release is control and visibility. Three pieces stand out:

Explicit breakpoints. You can now mark exactly which prefix of a prompt should be treated as cacheable — system instructions, tool definitions, a long document you're grounding on — instead of relying on the API to infer the boundary implicitly. That matters because implicit caching is fragile: append one token in the wrong place and the whole prefix invalidates.

A diagnostics tool. When a request misses cache, OpenAI now tells you why — tools_changed, a model swap, a settings modification — along with the estimated token cost of that miss. Previously you'd see a bill and have to guess. Now you get a reason code.

A dashboard. Aggregate cache hit rate and input composition, visible per project. This is the kind of thing you only miss once you've had it — before this, teams built their own logging just to answer "why did our cache hit rate drop this week."

On top of that, GPT-6 lets you change reasoning effort mid-conversation via a configuration_update append rather than a request-level parameter change, specifically so that adjusting effort doesn't blow away your cached prefix. That's a small detail, but it tells you what OpenAI is actually optimizing for: long-lived, stateful sessions where the model's settings evolve but its grounding context doesn't.

Why this is an infrastructure story, not a model story

It's tempting to skim past caching updates because they don't move any capability benchmark. But cost and latency are capability constraints in practice. An agent that reprocesses its full system prompt, tool schema, and conversation history on every turn is bounded by that reprocessing cost long before it's bounded by the model's reasoning. Caching is what makes persistent agents — the kind that stay open for hours doing code review, research, or multi-step tool use — economically viable at all. A 90% discount on the token volume that dominates a long session (the accumulated history) changes the shape of what you can build, even though it changes nothing about what the model can think.

The diagnostics piece matters for the same reason. Cache invalidation has been a silent tax: teams would restructure a system prompt, add a tool, or bump a setting, and watch costs creep up without an obvious cause. Naming the miss reason turns a mystery into a lint error.

Here's the shape of the change:

Diagram of GPT-6 prompt caching showing a shared prefix cached at a breakpoint versus a fresh tail, contrasting cache hits within a 30-minute window at up to 90% discount against cache misses from changed tools, model, or settings, with a diagnostics dashboard tracking hit rate

What I'd change in a production system

A few practical takeaways if you're running GPT-6 in production:

  • Put stable content first, volatile content last. Cache breakpoints only help if everything before the breakpoint is actually stable. System instructions and tool schemas belong at the front of the prompt; anything that changes per-request belongs at the end.
  • Treat tool schema changes as a cost event, not just a behavior change. Every time you edit a tool definition, you're invalidating cache for every session using it. The diagnostics tool now makes that cost visible — use it before you ship a schema tweak, not after the bill arrives.
  • Watch the 30-minute window. If your agent's idle gaps regularly exceed that, you're not benefiting from caching the way the pricing model assumes. That's a scheduling problem as much as a model problem — batch or debounce requests to stay inside the window where it's cheap to do so.
  • Use the dashboard as a design signal, not just a cost report. A dropping hit rate usually means something in your prompt construction got less disciplined. It's an early warning for prompt sprawl.

None of this required OpenAI to make GPT-6 smarter. It required them to make the system around GPT-6 more legible. That's a less exciting kind of release, but for anyone running agents in production, it's the kind that actually shows up in the invoice.

References
  1. 01Better prompt caching for GPT-6 — OpenAI