The 1-hour cache fix for batches has a break-even you cannot see
Anthropic's batch docs tell you to swap the 5-minute prompt cache for the 1-hour one. The swap costs a 60 percent heavier write on every miss, and the hit rate it has to reach to pay for itself is 57.6 percent at the bottom of Anthropic's own published band and outside that band at the top. A correction to my own September post.
On this page
Quick answer
As of 4 October 2026, Anthropic's batch documentation tells you to reach for the 1-hour prompt cache because batches outlive the 5-minute default. That advice is correct, it is unconditional, and the condition it leaves out is the entire decision. A 1-hour cache write is billed at 2 times base input where a 5-minute write is billed at 1.25 times, so the longer lifetime only pays if it lifts your cache hit rate by enough to cover a 60 percent heavier write on every miss. I worked the threshold out from the published multipliers. Against a 30 percent hit rate on the 5-minute cache you need 57.6 percent on the 1-hour cache just to break even. Against 98 percent you need 98.8 percent, which sits outside the 30 to 98 percent band Anthropic itself publishes for batch cache hits. And if switching the lifetime does not move your hit rate at all, the 1-hour cache is 12 to 16 percent dearer than the thing it was meant to fix.
This post corrects my own September write-up, which called that 2x write a one time cost.
The moment
I published a batch API field log on 16 September and left one sentence in it that has been quietly bothering me ever since. Describing the documented fix for a prompt cache that expires mid-batch, I wrote that the heavier write is a one time cost that the rest of the batch then reads at 0.1x instead of re writing at 1.25x.
In the same post, under What I did not verify, I wrote that I had run no controlled cost comparison between a batch with default caching and the same batch with a 1-hour lifetime, and that the resulting number was the one I actually wanted.
I still cannot run that batch. What I can do, and had simply not done, is the arithmetic over the multipliers Anthropic already publishes. It takes about four lines. It changed my answer.
Everything quoted below is read from Anthropic's batch processing reference and its prompt caching reference, both fetched on 4 October 2026. I have been careful to separate what the documents say from what I am deriving.
Finding 1: step two of the documented recipe is an action you cannot take
The batch page has a short section headed Using prompt caching with Message Batches. It opens by telling you that hits are best effort, because batch requests are processed asynchronously and concurrently, cache hits are provided on a best-effort basis, and that Users typically experience cache hit rates ranging from 30% to 98%, depending on their traffic patterns.
Then it gives three steps to maximise your hit rate. Step two, verbatim:
Anthropic's batch documentation, step 2 of 3: Maintain a steady stream of requests to prevent cache entries from expiring after their 5-minute lifetime.
That is good advice on the synchronous path, where you own the clock. Inside a batch you do not own the clock. You hand the platform a file of requests once, and the same page has just told you those requests run concurrently and in whatever order the platform chooses. There is no knob, no pacing parameter and no submission order guarantee that lets a caller space requests out inside a batch. Step two asks for something the API is specifically designed not to give you.
There is a charitable reading, and I want to put it up rather than skip past it: perhaps step two means maintain a steady stream of batches, which you genuinely can do. The page never says that, and step one scopes the whole list inward by asking for identical cache blocks in every Message request within your batch, so the natural reading is requests rather than batches. I think it is a sentence carried over from the synchronous caching guidance without being re-checked against the surface it now sits on.
Two smaller things sharpen it. First, step two's premise is the 5-minute lifetime, and roughly a thousand lines earlier the same page has a tip box telling you the opposite: Because batches can take longer than 5 minutes to process, consider using the 1-hour duration instead. One page, two positions on whether five minutes is enough.
Second, the worked example directly underneath the three steps ships the 5-minute default. Every cache block in it reads like this, with no lifetime specified:
{
"type": "text",
"text": "your large shared system prompt",
"cache_control": { "type": "ephemeral" }
}
The page then closes the example by saying those blocks are there to increase the likelihood of cache hits. So the one worked demonstration of batch caching in the documentation uses the lifetime that the same document elsewhere says is too short for a batch.
Finding 2: the break-even, and it is not the break-even everyone computes
There is a healthy body of writing on Claude's cache economics, and it is worth saying clearly that it mostly answers a different question from the one a batch author faces.
Ryan Skidmore's cache tokenomics piece derives a neat 62.5-minute rule for when to refresh a cache versus let it lapse. Falak Mahmood's write, read and TTL math is the most thorough treatment of the batch caching surface I found, and it already carries the 30 to 98 percent band, the stacking of the two discounts, and the rejection of pre-warming inside a batch. Both frame the break-even the standard way: how many reads does a write need before caching beats not caching.
Neither question is available to you in a batch. You cannot time refreshes, because you do not control when requests run. And you are not choosing between caching and not caching, because the three-step list has already told you to put identical cache blocks on every request. The only live choice is which lifetime to put in them.
So the useful question is narrower: at a given hit rate, which lifetime costs less. The multipliers are published. From the caching reference, 5-minute cache write tokens are 1.25 times the base input tokens price, 1-hour cache write tokens are 2 times the base input tokens price, and Cache read tokens are 0.1 times the base input tokens price.
The step that matters is noticing what a miss costs. A request carrying a cache block either hits, and is billed at the read rate, or misses, and writes, and is billed at the write rate. A miss is not a plain input charge. Hal Beri's work on cache savings and hit rate puts this well for the agent loop, noting a broken cache can land 24% more than not caching at all. In a batch the same mechanic applies, with a heavier write if you took the documented advice.
So for a shared prefix, the expected cost per token per request is the hit rate times the read multiplier plus the miss rate times the write multiplier. Set the two lifetimes equal and solve for the 1-hour hit rate:
# required 1h hit rate for the 1-hour lifetime to be no dearer than the 5-minute one
# h * read + (1 - h) * write is the expected multiple of base input per request
def required_1h_hit_rate(h5, read=0.10, w5=1.25, w1=2.00):
return ((w1 - w5) + h5 * (w5 - read)) / (w1 - read)
for h5 in (0.30, 0.50, 0.70, 0.90, 0.98):
print(h5, round(required_1h_hit_rate(h5), 3))
The batch discount halves both strategies identically, so it cancels and does not enter the comparison at all. Here is what it gives:
Scroll to see more
| 5-minute hit rate | 1-hour hit rate needed to break even | uplift required |
|---|---|---|
| 30% | 57.6% | 27.6 points |
| 50% | 69.7% | 19.7 points |
| 70% | 81.8% | 11.8 points |
| 90% | 93.9% | 3.9 points |
| 98% | 98.8% | 0.8 points |
Two things fall out, and the second one is the finding.
At the bottom of Anthropic's published band the uplift required is large. Going from a 30 percent hit rate to a 50 percent one sounds like a clear win and is not: on Sonnet 5.5 inside a batch that is 1.05 dollars per million prefix tokens against 0.905, which is 16 percent worse. You need 57.6 percent before you are level.
At the top of the band the requirement leaves the band. If your 5-minute hit rate is already 98 percent, the figure Anthropic names as the ceiling of normal experience, the 1-hour lifetime has to deliver 98.8 percent to break even. There is no headroom inside the published range for it to find. Switching lifetime at that end of the band is a rate rise, not an optimisation, and holding the hit rate flat at 98 percent costs 12.2 percent more.
The band is published without saying which lifetime it describes, and I am not going to pretend to know. What I can say is that the relationship is robust to that: the requirement is 57.6 percent at the band floor and 98.8 percent at its ceiling whichever lifetime the band was measured on. It is also robust across models. The read multiplier varies, 0.05 times base input on Claude Opus 5.5 and 0.025 on Claude Fable 5.1 and Claude Mythos 5.1, and that moves the floor only from 57.6 to 56.6 percent and leaves the ceiling at 98.8.
Where this leaves my September post. That post's reasoning holds if the 1-hour lifetime delivers near total hits. At a 100 percent hit rate it is about 9 times cheaper than a 30 percent 5-minute cache, which is a rout, and the enthusiasm was earned. What it did not do was state the condition. In a surface the vendor describes as best effort, with a published floor of 30 percent, assuming near total hits is the whole argument rather than a detail of it. Right mechanism, unstated condition.
Finding 3: the one pre-warming route still open, and a constraint neither page states
Pre-warming is the obvious way out of all of this: write the entry before the batch, then let the batch read it. The caching reference documents it as a request with no output budget, and the batch page bans it inside a batch, because an ephemeral cache entry written during batch processing would likely expire before the follow-up request runs. Falak Mahmood's piece covers that ban, and I covered it in September as a second instance of the same clock.
What neither page says is that the ban is scoped to requests inside the batch. A pre-warm on the synchronous path, asking for the 1-hour lifetime, followed by a batch submission, is not prohibited by anything I can find. That is the shape I would try first.
It carries a constraint that is derived rather than stated, and it comes from joining two sentences that live on different pages. The caching reference says Caches are isolated per workspace, ensuring data separation between workspaces within the same organization. The batch page says Batches are scoped to a Workspace. So a pre-warm has to be issued by a key in the same workspace the batch will run in, or it writes an entry into a cache the batch cannot reach. Workspace-scoped keys make that an easy mistake to make and a silent one to pay for, because the failure mode is not an error, it is the write premium twice. That is the same shape as the advisor tool, where both documented ways of measuring your token spend under-report it in the same direction: a cost surface that reads as working is the expensive kind of wrong.
I have not tested this. It is a derivation from two documented facts, and it is the first thing I would instrument rather than the thing I would ship.
What I did not verify
This is a documentation reading plus arithmetic over published prices. On a post about cost that distinction matters more than usual.
What I read directly: Anthropic's batch processing and prompt caching references on 4 October 2026, and the five third-party pages I credit above, all fetched and read rather than skimmed from search snippets.
What I did not measure: any actual cache hit rate. I did not submit a batch with the 5-minute default and the same batch with the 1-hour lifetime and compare invoices. That experiment is still the one I want, and it is still the one I have not run. Everything in finding 2 is arithmetic over multipliers, which is a strong form of argument about relative cost and no form of argument at all about what hit rate either lifetime actually achieves in your traffic. The threshold tells you what to measure, not what you will find.
What I could not read: there is a thread on r/ClaudeAI dated 6 September 2026 asking precisely when the 1-hour cache beats the default. It is behind a login wall and two separate browsers would not get me in. If somebody in there has already derived this same threshold, I am not claiming to have got there first.
What I deliberately did not claim: anything about how other vendors price the equivalent choice. I have not read their current reference documentation, and writing a cross-vendor comparison from memory is how wrong numbers get published.
What may be wrong in my model: I assume every request in the batch carries the cache block, so a miss always writes. If the platform coalesces concurrent writes for an identical prefix and bills one of them, the floor requirement drops. Nothing I read says either way.
Postscript: I have now corrected a sentence I wrote about a cache that lives for five minutes, eighteen days after writing it, which is roughly five thousand cache lifetimes of reflection for one subordinate clause.
Written by
M. PatelFrequently asked questions
Does the 1-hour prompt cache always save money in a Message Batches request?
No. A 1-hour cache write is billed at 2 times base input and a 5-minute write at 1.25 times, so the longer lifetime only pays if it raises your cache hit rate enough to cover the heavier write on every miss. Against a 30 percent hit rate on the 5-minute cache you need 57.6 percent on the 1-hour cache to break even, and if the hit rate does not move at all the 1-hour cache is 12 to 16 percent dearer. Figures read from Anthropic's prompt caching reference on 4 October 2026.
What cache hit rate does the 1-hour lifetime need to beat the 5-minute default?
It depends on what the 5-minute cache is already achieving. Required 1-hour hit rates are 57.6 percent against a 30 percent baseline, 69.7 percent against 50 percent, 81.8 percent against 70 percent, 93.9 percent against 90 percent and 98.8 percent against 98 percent. The Message Batches discount halves both options identically, so it cancels out of the comparison entirely.
Why can I not simply follow Anthropic's three steps for maximising batch cache hits?
Step two asks you to maintain a steady stream of requests so entries do not expire after their 5-minute lifetime. Inside a batch you submit once and the platform runs the requests concurrently and in an order it chooses, so there is no pacing knob available to a caller. The step reads as synchronous caching guidance that was not re-checked against the batch surface.
Can I pre-warm a prompt cache before submitting a Message Batches job?
A zero output budget pre-warm request is rejected inside a batch, because an entry written during batch processing would likely expire before the follow-up request runs. The ban is scoped to requests inside the batch, so a synchronous pre-warm asking for the 1-hour lifetime before submitting looks permitted. Note that caches are isolated per workspace and batches are scoped to a workspace, so the pre-warm must use a key in the same workspace. That route is derived from the documentation and is untested.
Is a prompt cache miss inside a batch just charged at the normal input rate?
No, and this is the step that makes the arithmetic counterintuitive. A request carrying a cache block either hits, billed at 0.1 times base input, or misses and writes, billed at the write multiplier for its lifetime. A cache you never read is therefore a surcharge on the part of the prompt you were trying to make cheap, not a neutral outcome.
Keep reading
Anthropic Batch API in 2026: the constraints nobody mentions until you hit them
The Batch API's default prompt cache expires in 5 minutes while the batch itself runs for up to an hour, so the two cost levers quietly cancel each other out. Six documented constraints, and the single assumption underneath all of them. Field log, September 2026.
Claude's token counter stops at the advisor's first call
The advisor tool is the only server tool Anthropic's count_tokens endpoint accepts, and the count covers the executor's first sampling call only. The top-level usage object is executor-only too, so both default meters miss the advisor sub-inference. Arithmetic on Anthropic's own published example puts 75.2 percent of output tokens outside the headline figures.
Continuing a context window truncation returns a 400
Two stop reasons mean your response was cut off, and the continuation helper on Anthropic's own stop reasons page only handles one of them. model_context_window_exceeded falls out of the loop and is returned as complete. Adding it to the condition is worse: the continuation re-sends a full window as input, which is documented as a 400 on every model.