New-ZZZ
RU / EN
Developer Tools 22 September 2026

OpenAI Boosts GPT-6 Caching for Faster, Cheaper AI Agents

N
New-ZZZ desk
OpenAI Blog · 6 days ago

OpenAI has announced a significant overhaul of its prompt caching system for the GPT-6 model family, a major upgrade designed specifically to enhance the performance, reliability, and cost-efficiency of persistent AI agents. These advanced agents are capable of handling complex, multi-step tasks—ranging from comprehensive code refactoring across entire codebases to generating highly detailed, well-researched documents and professional presentations. The core challenge addressed by this update is the nature of agent workflows: they are not single, isolated requests, but rather a continuous series of API calls that build upon each other. Crucially, these sequential calls frequently carry forward identical or highly similar instructions, tool definitions, and contextual information from previous turns. Historically, OpenAI utilized caching to store and reuse this shared context, which drastically reduced response times and offered developers substantial cost savings, sometimes up to 90% discounts on the input tokens that were successfully cached. The new system builds upon this foundation by delivering higher cache hit rates by default, and it introduces a powerful financial incentive: cache discounts are now provided for eligible shared prefixes that are reused within a 30-minute time window. This focus on maximizing cache hits and providing granular cost incentives fundamentally changes how developers can architect long-running, expensive AI applications.

To empower developers to fully utilize these advanced caching capabilities, OpenAI has launched several sophisticated monitoring and diagnostic tools. The new Prompt Caching Dashboard is a central hub that provides developers with deep visibility into how much of their application's input is being served directly from the cache. Users can track hit rates over extended periods and utilize an input composition chart. This chart is invaluable because it allows developers to visually compare the proportion of tokens that were cached versus those that were uncached. This level of granular tracking enables teams to proactively spot any unexpected drops in cache hit rates, allowing them to evaluate precisely how changes made to their application's logic or inputs might negatively impact the overall caching performance. Furthermore, when a developer encounters an unexpected cache miss—a failure to reuse cached context—the dedicated prompt caching diagnostics tool steps in. This tool is designed to be highly forensic, enabling the user to compare the failing request against a recent, successful response. By doing so, the developer can pinpoint the exact cause of the failure, identifying whether the issue stems from changes in the model itself, modifications to the available tools, changes in settings, or subtle alterations in the input prompt. The system even provides an estimate of the number of affected tokens, giving the developer a clear measure of the potential impact and guiding them on the necessary optimizations to maximize future cache reuse.

Beyond mere monitoring, the update provides developers with unprecedented control over the caching mechanism. A key feature is the introduction of explicit cache breakpoints. These breakpoints allow developers to make conscious, architectural decisions about exactly which prompt prefixes they want the model to reuse across multiple requests. The accompanying guide details how to implement these breakpoints, how long the cached prefixes remain eligible for reuse, and, critically, how changes to the available tools or the input data might affect the longevity and usability of the cached context. This level of control moves caching from a passive benefit to an active, programmable component of the application architecture. Furthermore, OpenAI has addressed the common scenario where an agent needs to adjust its internal thinking process without losing its accumulated context. On the GPT-6 models, developers can now adjust the 'reasoning effort' between responses. This means a developer can decide to raise the effort level for a particularly difficult task—requiring deeper, more complex thought—or lower it for a simple follow-up query. This adjustment is achieved by appending a specific configuration_update while leaving the request-level reasoning effort untouched, thereby allowing the model to adapt its cognitive load while preserving the valuable, reusable context stored in the cache. This ability to decouple reasoning effort from context preservation is a massive leap forward for building sophisticated, stateful agents.

To ensure the cache remains robust even as the agent's operational requirements evolve, OpenAI has provided best practices for preserving cache integrity when tools and instructions change. As an agent's tool use needs change—for instance, if it starts needing a new API or if the schema of an existing tool is updated—the developer must take proactive steps to keep the tool definitions, schemas, and the overall ordering stable. Instead of removing tool definitions entirely, which would break the cache, developers should utilize the allowed_tools parameter to restrict the model to only the tools currently relevant for the task, or set tool_choice to none when no tools are needed. For managing evolving instructions, the system recommends using new developer messages appended towards the end of the context. This technique allows developers to introduce new, overriding instructions without invalidating the older, foundational instructions that are still necessary for the agent's core functionality. Finally, to minimize the perceived latency for the end-user, OpenAI introduced the concept of 'prewarming.' Prewarming allows an application to proactively prepare and load known context—such as shared instructions, essential tool definitions, or foundational reference material—during the application's startup phase, before the user even submits their first query. By moving this necessary processing time out of the user's waiting window, the model can begin responding much sooner when the actual request arrives, significantly improving the perceived speed and responsiveness of the entire system. These optional, yet powerful, controls allow developers to fine-tune the caching mechanism to perfectly match the unique demands and workload patterns of their specific AI applications, moving beyond the engine's default performance to achieve optimal, tailored efficiency.

Why it matters

  • —It dramatically reduces the operational cost of running complex, multi-step AI agents by maximizing context reuse.
  • —It provides developers with advanced diagnostic tools to pinpoint and fix cache failures, improving reliability.
  • —It introduces fine-grained control over context preservation, allowing agents to adapt their reasoning effort and toolsets without losing accumulated memory.

Key facts

  • The new system offers higher cache hit rates by default and provides discounts for shared prefixes reused within 30 minutes.
  • Developers gain access to a Prompt Caching Dashboard and a diagnostics tool to monitor and troubleshoot cache performance.
  • New controls allow adjusting the model's 'reasoning effort' between responses without breaking the reusable context.
  • The ability to 'prewarm' the cache means necessary context can be loaded during startup, reducing user-perceived latency.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related