Dani Reyes10 min read3 views

Your Claude conversation has two prefixes, and only one of them errors

A thinking block stays valid against the bytes you sent. The prompt cache keys on the prompt the model renders. The rows where those two disagree cost money with no error.

Flat schematic on deep navy. Two rows of blocks share the same change point. The upper row stays lime green past it, showing thinking blocks still valid. The lower row turns orange from that point on, showing the prompt cache lost from there. No text.
Flat schematic on deep navy. Two rows of blocks share the same change point. The upper row stays lime green past it, showing thinking blocks still valid. The lower row turns orange from that point on, showing the prompt cache lost from there. No text.
On this page

Quick answer

Written 5 October 2026, against the Claude docs as they read today.

A Claude conversation carries two prefix checks, not one. They are computed over different things, and they disagree.

The thinking-binding check compares what you sent: the top-level system prompt, the tools array, and every message before the block. The prompt cache keys on the prompt the model actually renders, which includes configuration you never wrote into messages at all.

So a single change can leave every thinking block valid and still throw your entire cached prefix away.

One of those failures returns a 400 and you fix it the same afternoon. The other one just bills you, quietly, forever.

The moment

I spent a morning this week chasing a 400 on a long tool-use loop.

The error named an invalid thinking block. Standard stuff. I had rebuilt the message array between turns, which is the documented way to break the binding.

What stopped me was the fix. I made the loop append-only, the 400 went away, and my cache reads did not come back.

That should not happen. If the prefix is intact enough to keep a signed thinking block valid, why is the cache still cold?

Because those are two different prefixes. I had been reading one table and paying for the other.

I had spent the morning treating the 400 as the whole problem. The 400 is just the half of the problem that shouts.

The preserved thinking page has a table called what counts as an edit. The prompt caching page has a table called what invalidates the cache. They are on different pages, they use different column headings, and nobody prints them side by side.

I printed them side by side. The interesting rows are the ones where they disagree.

Finding 1: the two checks read different objects

The binding check is byte-scoped to what your client transmitted.

Anthropic, verbatim, on the thinking page: A block stays valid only while the top-level system prompt, the tools, and the messages before it are unchanged.

Three inputs. That is the whole prefix. Anything outside those three is invisible to the check.

The cache is scoped differently.

Anthropic, verbatim, on the prompt caching page: Cache keys are generated using a cryptographic hash of the prompts up to the cache control point.

Note the word prompts, not request fields. The rendered prompt is not the JSON you sent. Configuration gets compiled into it.

The thinking configuration goes in. The resolved effort level goes in. The speed setting goes in. Toggling web search rewrites the system prompt before the model sees it.

None of that is in system, tools, or messages. All of it is in the hash.

The cache also has a shape the binding check does not. It is layered, and the layers fall in order.

Anthropic, verbatim: Changes at each level invalidate that level and all subsequent levels.

The order is tools, then system, then messages. So a change high up takes everything below it with it, which is why a single added tool costs more than it looks like it should.

And the match is not fuzzy.

Anthropic, verbatim: Cache hits require 100% identical prompt segments, including all text and images up to and including the block marked with cache control.

100% identical. There is no partial credit, no nearest-prefix fallback. You either reproduce the rendered bytes exactly or you pay to write them again.

That is the whole asymmetry, and the docs state the reason for it exactly once, parenthetically, on a row about something else. The binding table annotates server-side compaction like this:

Anthropic, verbatim: Valid (the check compares what you sent, not the server's edited copy)

Compares what you sent. Hold that sentence against a cache key built from the rendered prompt and the rest follows.

Finding 2: the vendor joins the tables for exactly one row

Credit where it is due. Anthropic does state the join, once, for the effort parameter.

Anthropic, verbatim: Changing top-level output_config.effort between requests doesn't invalidate thinking, because effort isn't part of the prefix. Changing top-level effort does restart the prompt cache.

Two sentences, opposite verdicts, same edit. This is the clearest statement of the asymmetry anywhere in the docs, and it is buried in a section about a workaround.

I went looking for who else had written this down. One competitor, a September 2026 piece on prompt caching, reproduces the cache invalidation table in full and correctly: tool definitions, model, system prompt text, web search and citations toggles, speed setting, tool_choice, images, thinking and effort settings.

It is a good table. It contains zero occurrences of preserved thinking, prefix check, or thinking block.

That is the pattern across everything I read. The cache half is published. The binding half is published. The join is not.

Finding 3: one mechanism predicts three more rows

Here is why the join matters. If the cache sees rendered configuration and the binding check sees only your three inputs, then every server-side prompt rewrite is cache-visible and binding-invisible.

That is not one row. It is a class.

Scroll to see more

What you changeThinking blocksCache
Append messages at the endvalidkept
Top-level effortvalidmessages lost, tools and system model-specific
Web search togglevalidsystem and messages lost
Citations togglevalidsystem and messages lost
Speed setting, standard to fastvalidsystem and messages lost
tool_choicevalidmessages lost
Add, move or remove cache_control markersvalidbreakpoint moves, so reads move with it
Remove thinking blocks from the start of historyvalidprefix changes from that point
Edit system or toolsinvalidlost

The first column for rows two through eight is sourced from the binding table, which rules that changing any request parameter outside system, tools, and messages leaves later thinking blocks valid, and separately lists cache_control marker moves as valid.

The second column is the cache table. For the web search row, verbatim: Enabling/disabling web search modifies the system prompt. You did not touch system. The renderer did. The binding check never sees it; the hash does.

Those middle rows are the dangerous ones. They are silent in both directions. No 400, no warning, no dropped block in the response. Just a cold cache and a bigger invoice.

Why is this easy to miss? The two tables do not share a vocabulary. One asks whether later thinking blocks are valid or invalid. The other asks whether the tools, system, and messages caches are kept or lost. Neither word appears in the other table. If you are debugging a 400 you read the first one and stop, because it answered your question.

The detection trick, if you want one, is that these failures have a fingerprint in the usage counters. A silent prefix change shows up as cache_read_input_tokens near zero on a request you expected to be warm, with cache_creation_input_tokens close to the full conversation size. The thinking blocks come back fine. Nothing in the response says anything is wrong, because by the API's own rules nothing is.

The row I would stare at longest is removing thinking blocks from the front of the history. The binding table rules it valid and annotates the cost as the model losing that reasoning. True, and incomplete. Removing content from the start of messages changes the prefix from position zero, so the cache cost is total. The named cost is the smaller one.

Finding 4: every safe escape is append-only

The useful half of this is that the escapes all have the same shape.

Where the docs give you a way to make a change without editing the prefix, it is always append-only, and it is always safe on both checks at once.

New instructions go in a role: "system" message appended to messages, not into the top-level system field. Tool changes go in a tool_addition or tool_removal block. Effort changes go in a per-message output_config:

json
{ "role": "system", "content": [], "output_config": { "effort": "low" } }

That one needs the mid-conversation-output-config-2026-07-01 beta header. Trimming goes to server-side compaction or context editing rather than client-side slicing, because the check compares the conversation as you sent it.

So the rule I am taking away is simpler than the table. Do not mutate the prefix; extend it. Both checks reward the same discipline, which is presumably the point.

It also tells you where to look when something goes wrong. If you get a 400, you edited system, tools, or an earlier message. If you get no error and no cache reads, you changed something that renders into the prompt without touching those three. Those are different bugs with different fixes, and the response tells you which one you have only if you know there are two.

One asymmetry does run the other way, and it is worth knowing. A thinking block the API drops is a cache event even though you did nothing wrong.

Anthropic, verbatim, on the thinking page: A thinking block the API drops under any preserved-thinking condition changes the cached prefix from that block's position onward. Blocks passed back unchanged keep the cache intact.

Pair that with how a model mismatch is handled, verbatim: The API drops a block the current model can't read, without an error and without billing it.

Without an error. So routing one turn of a conversation to a model that cannot read the previous model's thinking drops the block silently and moves your cached prefix silently. Two invisible effects from one routing decision.

What I did not verify

I have no API key on this machine, so none of this is measured. Everything above is read off the current documentation and joined. I did not observe a single cache_read_input_tokens value.

Specifically unverified:

The web search, citations, and speed rows in my table are predictions. The cache column is documented. The thinking column is my inference from the binding table's rule about parameters outside system, tools, and messages. I did not find either page stating those three joins directly, and I did not test them.

The cache_control row is the weakest. I know marker moves keep thinking valid, because the binding table says so. What a moved breakpoint does to existing cache entries is not something I can state cleanly from the docs.

One row I deliberately left out. The binding table rules that a rotating signed URL returning the same bytes keeps thinking valid, which tells you the binding check compares fetched bytes rather than the URL string. Whether the cache hash keys on the bytes or on the literal URL is the difference between a warm cache and a guaranteed miss on every request, and I could not resolve it from either page. If you serve images from presigned storage URLs, that is worth measuring before you trust it.

I also did not read Google or OpenAI reference docs for this, so nothing here is a cross-vendor claim. The competitor piece I cite is accurate on the cache table as of September 2026; I am crediting it, not auditing the rest of it.

The honest test for all of this is two requests that differ only in the field under test, and a diff of the three usage counters. I could not run it today.

Postscript: I fixed the 400 in an hour and the cold cache is still sitting there, which is roughly the correct ratio of attention to cost in my work generally.

D

Written by

Dani Reyes

Frequently asked questions

Does changing effort invalidate Claude thinking blocks?

No. Anthropic states that changing top-level output_config.effort does not invalidate thinking, because effort is not part of the prefix. It does restart the prompt cache, so the thinking blocks survive and the cached prefix does not.

Why did my prompt cache go cold without a 400 error?

The thinking-binding check only reads the top-level system prompt, the tools array, and the messages before the block. The cache keys on the rendered prompt, which also includes thinking configuration, effort, the speed setting, and toggles such as web search that rewrite the system prompt server-side. Changing one of those is invisible to the binding check and visible to the cache.

How do I change instructions mid-conversation without breaking either one?

Keep the prefix append-only. Append a role system message to messages instead of editing the top-level system field, use tool_addition or tool_removal blocks for tool changes, carry effort changes in a per-message output_config, and trim with server-side compaction rather than slicing the array on the client.

The 1-hour cache fix for batches has a break-even you cannot see

Anthropic's batch docs tell you to swap the 5-minute prompt cache for the 1-hour one. The swap costs a 60 percent heavier write on every miss, and the hit rate it has to reach to pay for itself is 57.6 percent at the bottom of Anthropic's own published band and outside that band at the top. A correction to my own September post.

12 min read31

Your JSON schema is a second cached artifact

Structured outputs and strict tool use turn your JSON schema into a separate cached artifact. It has its own 24 hour lifetime, invalidation rules that invert what you would guess, and a retention boundary that zero data retention does not cover.

10 min read84

Continuing a context window truncation returns a 400

Two stop reasons mean your response was cut off, and the continuation helper on Anthropic's own stop reasons page only handles one of them. model_context_window_exceeded falls out of the loop and is returned as complete. Adding it to the condition is worse: the continuation re-sends a full window as input, which is documented as a 400 on every model.

11 min read24