Dani Reyes11 min read6 views

Continuing a context window truncation returns a 400

Two stop reasons mean your response was cut off, and the continuation helper on Anthropic's own stop reasons page only handles one of them. model_context_window_exceeded falls out of the loop and is returned as complete. Adding it to the condition is worse: the continuation re-sends a full window as input, which is documented as a 400 on every model.

Two horizontal bars against a vertical line marking the context window ceiling. The top bar shows input in slate and generated output in lime, ending exactly at the line. An arrow folds that whole bar down into the second bar, where all of it is now input, past the ceiling.
Two horizontal bars against a vertical line marking the context window ceiling. The top bar shows input in slate and generated output in lime, ending exactly at the line. An arrow folds that whole bar down into the second bar, where all of it is now input, past the ceiling.
On this page

Quick answer

Checked on October 3, 2026. Two different stop_reason values mean your response was cut off, and the continuation helper published on Anthropic's own Anthropic stop reasons page only handles one of them. Its loop condition is a single equality test against max_tokens, so a response that stopped with model_context_window_exceeded falls straight out of the loop and is returned as though it were finished. The table at the top of that same page says that value means the response is truncated.

The obvious repair is to add the second value to the condition. That repair is worse than the bug. model_context_window_exceeded fires when generation has already filled the context window, so a continuation turn re-sends everything that just filled it, this time as input, plus a short instruction on top. Input alone over the window is documented as a 400 on every model. You turn a quiet truncation into a hard request failure, on the one path where you had no recourse to begin with.

The moment

I had a summarisation job over long legal documents. It had run for months. The pattern was the boring one: set max_tokens high so the summary is never clipped, check stop_reason, continue if it says max_tokens, hand the text to a parser.

Then a batch came back with summaries that stopped mid-clause. Not many. Maybe one in forty. Every one of them was HTTP 200. Every one of them had gone through my continuation helper, which is to say my continuation helper had looked at them and decided they were complete.

I spent an afternoon convinced I had a streaming bug, because the only thing I checked was whether the text looked finished and whether the status code was fine. Both of those are the wrong question. The field that had the answer was one I was reading and then throwing away, because my code compared it to exactly one value and treated everything else as success.

The documents had grown. The prompt template had not changed. Nobody made this change, which is the part that makes it hard to find.

Finding 1: The published helper breaks on the other truncation

Anthropic's stop reasons page carries a worked helper called get_complete_response. It is the thing you copy when you want long outputs to survive truncation. The loop is short, and the whole behaviour sits in one line.

python
# from Anthropic's stop reasons page
if response.stop_reason != "max_tokens":
    break

Anthropic, verbatim: if response.stop_reason != "max_tokens":

Six of the seven possible values take that branch. Two of the six are genuinely complete and nobody should mind. model_context_window_exceeded is not one of those two, and the quick reference table ten lines above the helper says so.

Anthropic, verbatim: "Treat the response as truncated."

So on a single page, the prose classifies this value as truncation and the sample code returns it as a finished answer. The helper is not wrong about max_tokens. It is just written against a two-state world, and the enum has had seven states for a while.

Both documents are on the stop reasons page.

Finding 2: It is the only row in the table that names no recovery

I went through the quick reference table row by row, because the shape of it turned out to be the finding rather than the contents.

Scroll to see more

valuewhat the table tells you to do
end_turnuse the response
max_tokensraise the limit, or continue
stop_sequenceread which sequence fired
tool_userun the tool, return the result
pause_turnsend the assistant content back
refusalread the details, retry on a fallback model
model_context_window_exceededtreat the response as truncated

Six rows name an action. One row names a classification. That asymmetry is not an oversight and I want to be clear that I think the table is right: there is no action to name. You cannot raise a limit you did not set, and the thing you would normally reach for, continuing the turn, is the subject of the next finding.

The defect is not in the table. It is that the code sample below the table does not implement the distinction the table just drew.

Finding 3: The obvious fix converts a 200 into a 400

This is the part I did not expect, and it is arithmetic rather than opinion.

Start with what triggers the value. The context windows page is explicit that there are two overflow states and only one of them raises.

Anthropic, verbatim: "If the input alone already exceeds the model's context window, the API returns a 400"

That is invalid_request_error, with the message "prompt is too long", on every model, and it is from the context windows page. The second state is the interesting one.

Anthropic, verbatim: "If generation then reaches the context window limit, it stops with" model_context_window_exceeded.

Read those two together. The stop reason means generation ran until input plus output reached the window. Call the window W, the input I and the generated output O. At the moment the API hands you that response, I plus O is at the ceiling.

Now apply the continuation pattern. The helper builds its next request from the original prompt, the accumulated response, and an instruction.

Anthropic, verbatim: "Please continue from where you left off."

The new input is therefore I plus O plus that instruction. I plus O is already W. The instruction is a dozen or so tokens on top. The new input is strictly over the window, which is precisely the first overflow state, which is the 400.

Concretely, with a 200,000 token window: 190,000 tokens of input, a generous max_tokens, generation stops at 10,000 tokens of output with model_context_window_exceeded. The continuation request carries 190,000 plus 10,000 plus the instruction. That is 200,000 and change, against a 200,000 ceiling, and it fails before a single token is generated.

So the three available behaviours are: leave the helper alone and silently publish truncated text, add the value to the condition and get a 400 on the retry, or do the actual thing, which is to reduce the input. Only the third is a fix, and it is not a change to the loop. It is context editing, compaction, or dropping part of the conversation, and which part to drop is a product decision rather than an API one.

One honest caveat on the boundary. The doc says "exceeds", so at exactly W you are not over, and whether generation stops at W precisely or a few tokens short is not specified to the token. The continuation instruction is what reliably carries you across the line. If you append nothing at all and re-send, you land at or just under the ceiling with effectively zero room to generate, which is a 200 with no useful output rather than a 400. Neither outcome is a continuation.

Finding 4: The field is still writing four-value handlers

I wanted to know whether this was just me, so I read what else is published on handling stop reasons.

The most directly on-topic guide I found, dated September 17 2026 and updated the following day, opens its explanation like this.

ZeroLabs, verbatim: "Handling stop reasons correctly requires understanding the four distinct termination states returned by the API:"

The four are end_turn, max_tokens, stop_sequence and tool_use. Its TypeScript sample switches on those four and ends with a default branch that throws an "Unhandled stop_reason" error, which means a production system using it as published raises on three values the API can return. That page is here.

In fairness, and this matters: its examples pin claude-3-5-sonnet-20241022, and on models that old model_context_window_exceeded requires an opt-in beta header. Against its own example model the four-value list is defensible. It is the generality of the framing, on a page updated in late September, that does not survive contact with a current model.

I should also be clear about what is not novel here. Two of the pages I read enumerate all seven values properly, with handling notes for each: an agentic-loops reference and a Claude API error reference. The seven-value enum is not a secret and I am not claiming to have discovered it.

What I think is worth noticing is the classification. The three values the four-value model leaves out are pause_turn, refusal and model_context_window_exceeded, and those are exactly the three where the content in your hands is not usable as it stands. One needs a continuation, one must not be shown as an answer, and one is truncated. The four that get handled include both of the cleanly-finished cases. The omissions are not random, and they all fail in the same direction, which is to say silently.

I have written about the refusal path before, in a note on refusal fallback, and it has the same shape: a 200 that a naive handler files as an answer.

Finding 5: Where I disagree with the best page I found

The agentic-loops reference above is the strongest treatment of this I read. It has the full seven-value table, and it explicitly names ignoring pause_turn as an anti-pattern, which is more care than most pages take. Its advice for the value in question is one line.

That reference, verbatim: "Treat similarly to max_tokens"

and then

the same reference, verbatim: "Continue, summarize, or compact context."

The second and third options are right and are what I ended up doing. The first one is the one that cannot work, for the reason in Finding 3, and lumping this value in with max_tokens is the exact mental model that produces the broken repair. It is one word in an otherwise excellent table, and I would not mention it except that "treat it like max_tokens" is precisely the instinct that makes the 400 happen.

For a different framing of the two overflow states that I found genuinely clarifying, and which pushed me to go and check the arithmetic, there is a field note on context overflow whose description of the 200 case I have not improved on.

That note, verbatim: "It is a truncated answer wearing a success code"

Its remedy is a pre-flight token count before you send, which is the right place to spend the effort and is a different question from the one I am asking here. Mine is about what your code does after a response that already stopped.

What I did not verify

  • I did not reproduce the 400 against the live API. The claim in Finding 3 is a derivation from two documentation pages plus the helper's own construction of its next request, not an observed response. I believe it, I have not watched it happen, and the boundary case at exactly the window size is genuinely uncertain as described above.
  • I did not test the SDK typing claim. The stop reasons page notes this value is typed in the SDK beta namespace while the Messages API reference lists it in the stable enum. I read both and did not install anything to check which types a current stable SDK actually ships.
  • I did not check Bedrock or Vertex. Partner platforms diverge, and one of the search results I passed over suggested a different set of stop reason values on at least one of them.
  • I did not compare other providers. I wanted to know whether anyone else splits "your limit" from "the model's limit" into two signals, and I could not reach the relevant reference documentation to check. Rather than assert it from memory, I left it out.
  • I did not exercise the helper's other rough edge. It collects only the first text block from each response, so extended thinking or multi-block content would lose material for reasons unrelated to anything above.
  • I did not measure how common this is. One batch in forty is my number from one job, and it is a fact about my document sizes rather than about the API.

Postscript: the bug was never the missing text. It was that I had written a comparison against one value and then read the result as a boolean, and an enum that grows underneath a boolean does not raise anything. It just gets quieter.

D

Written by

Dani Reyes

Frequently asked questions

What does stop_reason model_context_window_exceeded mean?

Generation ran until input plus output reached the model's full context window, so the response is truncated. It is distinct from max_tokens, which means you hit the limit you set yourself. Anthropic's quick reference table tells you to treat the response as truncated, and unlike every other row it names no recovery action, because there is no limit of yours to raise.

Can I continue a response that stopped with model_context_window_exceeded?

No. The stop reason means input plus output already reached the window, so a continuation turn re-sends all of it as input plus an instruction on top. Anthropic's context windows page documents that input alone over the window returns a 400 invalid_request_error with the message prompt is too long, on every model. The repair is to reduce the input through context editing or compaction, not to continue the turn.

Why does the documented get_complete_response helper miss it?

Its loop condition compares stop_reason against the single value max_tokens and breaks on everything else. Six of the seven possible values take that branch, and model_context_window_exceeded is one of them, so a truncated response is returned as though it were finished. The helper is correct about max_tokens; it is written against a two-state world.

Which stop reasons does a four-value handler miss?

pause_turn, refusal and model_context_window_exceeded. Those are exactly the three where the content you hold is not usable as it stands: one needs a continuation, one must not be shown as an answer, and one is truncated. The four that commonly get handled include both of the cleanly finished cases, so the omissions all fail silently.

Claude's token counter stops at the advisor's first call

The advisor tool is the only server tool Anthropic's count_tokens endpoint accepts, and the count covers the executor's first sampling call only. The top-level usage object is executor-only too, so both default meters miss the advisor sub-inference. Arithmetic on Anthropic's own published example puts 75.2 percent of output tokens outside the headline figures.

10 min read37

Priority Tier capacity is not extra capacity

An Anthropic Priority Tier commitment is drawn down alongside your ordinary rate limit, not on top of it. Reading the service tiers page next to the pricing page and the models list turns up a 20x burndown spread, a single cache read rate that only holds because of the exclusion list, and three of four current models excluded.

10 min read45