Prompt caching reuses the stable opening part of a prompt across API requests instead of reprocessing it every time, which cuts latency and can cut input cost sharply on repeated calls. It pays off when you resend a large stable prefix such as system instructions, long documents, or growing conversation history within a short window. It saves almost nothing on one-off prompts, tiny prefixes, or prompts where the opening changes every request. Start with automatic caching on a stable prefix, keep the cached prefix byte-identical and ordered, and check cache creation and read token fields before assuming savings.
What is prompt caching, and when does it actually save you money?
Prompt caching reuses a stable prompt prefix across requests to cut latency and input cost. It wins on large repeated prefixes sent within the cache lifetime and saves little on one-off, tiny, or constantly changing prompts.
Published · Updated · Evidence-linked, not search-volume ranked.
Why this question is current
Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.
- prompt caching explained · Google Suggest · US; English · checked 2026-10-09T22:30:00Z
Observed 6 of 6 completions: the base query plus meaning, Claude-specific, in-LLMs explainer, and response-caching variants. A same-day explainer and cost cluster, not a volume or ranking claim. - what is prompt caching stories by date · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-10-09T22:35:00Z
Index reported 786 hits, with same-day items on 2026-10-09 including API gateway and coding-agent tooling stories. Proves active same-day builder discussion in adjacent agent infrastructure, not a volume ranking. - Claude prompt caching documentation · Claude Platform documentation · global official documentation · checked 2026-10-09T22:40:00Z
Official prompt-caching page documents automatic and explicit modes, 5-minute and 1-hour TTLs, per-model pricing, invalidation rules, and monitoring fields. Proves documented product surface for cost and latency claims.
Who this helps
- Developers calling LLM APIs with large repeated context
- Builders running multi-turn agents where history keeps growing
- Teams trying to lower input-token spend without shrinking context
the short answer
Prompt caching stores the processed form of a prompt prefix so later requests with the same opening can resume from it. The provider checks whether the prefix up to your cache breakpoint matches a recent entry. On a hit it reuses the cached work, which lowers time to first token and bills the reused tokens at a reduced hit rate. On a miss it processes the full prompt and writes the prefix to cache once the response begins.
The docs describe this as best for prompts with many examples, large background context, repetitive instructions, or long multi-turn conversations. The default lifetime is 5 minutes from the request that writes or reads the entry, and each reuse refreshes that window. Response generation time counts against the window, so a slow 4-minute response leaves about 1 minute to reuse the entry under the default TTL.
automatic versus explicit breakpoints
Automatic caching is the simplest start. You add one top-level cache control field and the system applies the breakpoint to the last cacheable block, moving it forward as the conversation grows. The docs show the cache point advancing each turn so growing message history stays cached without updating markers.
Explicit breakpoints give finer control. You place cache control on individual content blocks to pin sections that change at different rates, such as a stable system prompt plus a slowly changing document section. The two modes can combine, with the automatic breakpoint consuming one of the available breakpoint slots. If the last block already carries an explicit breakpoint with the same TTL, automatic caching is a no-op, and conflicting TTLs or exhausted slots return an error rather than caching silently.
when caching saves real money
Savings come from large stable prefixes sent repeatedly. A support agent that resends the same policy manual plus conversation history, a coding assistant that resends the same repository context, or a batch job that shares one long instruction block are the textbook cases. Cache hits are billed at a fraction of base input price: 0.1x on most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1 and Mythos 5.1, per the docs pricing table snapshot checked 2026-10-09.
Concrete snapshot from that table: Sonnet 5 lists 2 dollars per million base input tokens, 2.50 for 5-minute writes, 4 for 1-hour writes, and 0.20 for hits and refreshes. Opus 5.5 lists 4 base, 5 for 5-minute writes, 8 for 1-hour writes, and 0.20 for hits. Haiku 4.5 lists 1 base, 1.25 for 5-minute writes, 2 for 1-hour writes, and 0.10 for hits. Writes cost more than base input, so the math only works when each written prefix earns multiple hits. Recheck the live pricing page before budgeting, because per-model numbers move.
when caching does almost nothing
One-off prompts gain nothing because there is no second request to hit the cache. Tiny prefixes gain little because there is little reused work relative to the write premium. Rapidly changing openings defeat the mechanism: a breakpoint placed on a block that changes every request, such as a timestamp or raw user input, writes a fresh entry each time and never hits.
Long gaps also defeat the default. If requests arrive more than 5 minutes apart, the entry expires before reuse unless you pay the higher 1-hour write rate. Batch workloads are explicitly best-effort on ordering, so the docs recommend writing the shared prefix to the 1-hour cache first and then submitting the batch.
setup rules that keep the cache hitting
Put the breakpoint at the end of a stable prefix, not on moving content. Keep the cached blocks byte-identical and in the same order across requests. Keep images and tool definitions stable, because adding or removing images or changing tool settings breaks the cache and forces a new write. The docs note a limited lookback window and minimum token thresholds, so check the current page for your model rather than assuming a short header will cache.
Separate stable from volatile content deliberately. A practical layout is system instructions first, then reference documents, then the cache breakpoint, then the changing user message and recent turns. That way the large stable part hits while only the tail is reprocessed.
monitoring and privacy notes
Measure before claiming savings. API responses expose cache creation and cache read input token fields, so compare those against base input tokens over a real session. A healthy setup shows creation once and reads many times. Repeated creation with near-zero reads means the prefix is shifting or expiring between calls.
On privacy, the docs describe cache keys as cryptographic hashes of the prefix, with caches isolated per workspace or per organization depending on platform, and never shared across organizations even for identical prompts. Treat that as the documented design, not as a reason to place secrets casually: keep credentials out of cached blocks wherever possible, and keep cached documents scoped to the workspace that needs them.
a useful next action
Pick one repeated call in your app or agent loop, enable automatic caching on it, and run a realistic session while recording creation versus read tokens. If reads dominate, keep the layout and consider the 1-hour TTL for slower loops. If creation dominates, the prefix is moving or expiring, so stabilize the ordering before spending more.
Pair caching with a spending control, covered in the related RepoRadar answer on hard API limits, because caching lowers unit cost but does not cap total usage.
Sources checked
- Claude Platform docs: Prompt caching ↗ checked · global official documentation
Automatic versus explicit breakpoints; 5-minute default and 1-hour TTL; per-model write and hit pricing with multipliers; prefix-match behavior and refresh-on-use; invalidation on images and tool changes; creation and read usage fields; hash-keyed per-workspace and per-organization isolation; batch best-effort behavior.
- Claude Platform docs: Pricing ↗ checked · global official documentation
Prompt caching framed as reusing processed prompt portions at a fraction of standard input price instead of reprocessing large system prompts, documents, or history on every request.
- Google Suggest for prompt caching explained ↗ checked · US; English
Same-day completions show explainer and Claude-specific intent around caching meaning and LLM use.
- HN same-day prompt caching search ↗ checked · global English-language developer community
High same-day index activity around prompt caching and adjacent agent infrastructure stories on 2026-10-08 and 2026-10-09.
RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.