Claude structured outputs can silently change your enum
Structured outputs promise you can stop validating. The same page documents four ways the output can miss your schema, and one of them returns a normal 200 with no stop_reason. Measured: the caveat says it covers strict tool use, and the strict tool use page never links it.
On this page
Quick answer
As of October 2026, Claude's structured outputs feature promises you will never need to validate its output again, and the same page documents four ways the output can fail to match your schema. Three of those four hand you a signal. The fourth is a silent change to the capitalization of your enum values, and the documented mitigation for it is to write exactly the defensive comparison the feature told you to delete.
That caveat says in terms that it applies to strict tool use too. It lives only on the JSON outputs page. The strict tool use page restates the guarantee six times and never links it.
The moment
I was replacing a hand-rolled retry loop. The old code called Claude, parsed JSON, validated against a schema, and retried on failure. The new code sets output_config.format and trusts the result, because that is what the feature is for.
The schema had one enum. A status field: In Progress, Needs Review, On Hold, Done. Four values, straight out of the ticket system they map to.
I was reading the page end to end before deleting the validator, which is the only reason I got to the section 2,500 lines down called ### Invalid outputs. I had already read the three bullets at the top that told me I did not need to.
Here is what I found, in the order it changed my mind.
Finding 1: the guarantee is stated three ways and then qualified four ways
The top of the structured outputs page is unambiguous. It says structured outputs "guarantee schema-compliant responses through constrained decoding" and then lists what you get:
Anthropic, verbatim: "Always valid: No more JSON.parse() errors", "Type safe: Guaranteed field types and required fields", "Reliable: No retries needed for schema violations"
That is the pitch, and for structural validity it is accurate. Constrained decoding really does make a syntax error impossible.
Much further down, a section opens with a sentence that has a different temperature:
Anthropic, verbatim: "While structured outputs guarantee schema compliance in most cases, there are scenarios where the output may not match your schema"
"In most cases" is doing a lot of work there, and the four scenarios it introduces are not equally visible to your code. Three of them tell you something went wrong:
- A refusal sets
stop_reasontorefusal, returns HTTP 200, and bills you for the tokens. The page is explicit that "the refusal message takes precedence over schema constraints". - A truncation sets
stop_reasontomax_tokensand the output may simply be incomplete. - An unsupported schema feature does not reach generation at all. The limitations list ends with "If you use an unsupported feature, you'll receive a 400 error with details."
Every one of those is a branch you can write. Check stop_reason, check the status code, done.
The fourth one is different in kind, and that difference is the whole post.
Finding 2: the one failure with no signal attached
The fourth scenario is headed Enum value casing, and it is worth reading slowly:
Anthropic, verbatim: "Structured outputs don't guarantee the capitalization of string enum and const values: Claude may return a value that differs from your schema only in capitalization, typically in the first letter of a word following a space."
So the drift is not random. It is specifically a capital letter appearing after a space, which means it needs a multi-word value to happen at all. The page's own example is an enum of Conversation Topic 1, Conversation Topic 2 and Conversation topic 3, where the output may come back as Conversation Topic 3 with a capital T that is not in your schema.
And then the sentence that separates this from the other three:
Anthropic, verbatim: "The response completes normally, with no error and no special stop_reason."
No error. No status code. No stop_reason to branch on. You get a 200, well-formed JSON, every required field present, correct types throughout, and one string value that is not a member of the set you declared.
The mitigation is one sentence later:
Anthropic, verbatim: "Compare enum values case-insensitively, and avoid enum values that differ only in capitalization."
Both halves of that are good advice and both halves cost you something the feature was sold on. "Compare case-insensitively" is a validation step on the output, which is the thing "No retries needed for schema violations" told you to stop writing. And "avoid enum values that differ only in capitalization" is a constraint on the schemas you are allowed to author, which is not in the list of unsupported features, because unlike everything on that list it does not give you a 400. It gives you a wrong answer.
The practical shape of this matters more than the wording. If your enum feeds a switch, an unmatched value falls to the default branch. If it feeds a database column with a check constraint, you get a write failure one layer away from the cause. If it feeds a Pydantic or Zod model, you do get an exception, which is the good case and is also only true because you kept the validator.
Finding 3: the caveat claims two features and sits on one page
This is the finding I would not have looked for if the first two had not sent me to the sibling page.
The caveat closes with Anthropic, verbatim: "This applies to both JSON outputs and strict tool use."
Strict tool use has its own page. On that page, measured today, the words capitalization, casing and case-insensitive occur zero times. What the page does carry is the guarantee, repeatedly:
Anthropic, verbatim: "Setting strict: true on a tool definition guarantees Claude's tool inputs match your JSON Schema"
Anthropic, verbatim: "Functions receive correctly-typed arguments every time", "No need to validate and retry tool calls"
And in a block literally headed Guarantees, Anthropic, verbatim: "Tool name is always valid (from provided tools or server tools)".
Now the part that is structural rather than rhetorical. The strict tool use page does send you to the JSON outputs page for schema caveats. It links the json-schema-limitations anchor twice. Across both pages that anchor is linked eight times. The invalid-outputs anchor, which is where the silent failure lives, is linked exactly once in the whole pair, and that one link is on the structured outputs page itself, inside the very accordion that enumerates the loud failures.
So a reader who follows the referral correctly lands on the list of things that return a 400, which is the list that cannot hurt them, and the paragraph about the thing that returns a 200 with the wrong value is one heading away and unlinked from the page that needs it.
I want to be fair about what this is. It is not a contradiction and nothing on either page is false. The caveat exists, it is written clearly, and it names its own scope accurately. It is a placement problem, and placement problems are the kind you only notice when you arrive from the other direction.
Finding 4: every enum in the vendor's own examples is immune to it
This is the measurement that explained to me why a documented, clearly written caveat is so easy to never hit.
The failure needs a string enum with a space in it. So I counted the enums in the two pages' worked examples.
The strict tool use page contains 12 JSON-style enum arrays. Ten of them are numeric, which the caveat does not touch at all, since it is scoped to string enum and const values. The remaining two are string enums, and both are the same value set: celsius and fahrenheit. Single words, lowercase, no spaces. Structurally incapable of exhibiting the drift.
Across both pages, the only multi-word string enum I could find is the Conversation Topic set inside the caveat itself. The example constructed to demonstrate the bug is the sole example in either document that could produce it.
The same held on the clearest third-party source I read. Timo Labs' study guide on this exact feature uses the enum goods, services, subscription, unclear, other. Five values, all single words, all lowercase, all immune.
None of this is a criticism of the examples. Short lowercase identifiers are good schema design and I would write them too. It is an explanation of the gap: you can follow every example on the page, ship, and never see this. Then someone models a real workflow state, writes In Progress and Needs Review because that is what the column already contains, and the enum goes from immune to exposed without anybody changing a line of configuration.
My rule out of this: if an enum value contains a space, either make it a single lowercase token and map to the display string in your own code, or keep a case-insensitive comparison. Not both guarantees and no validator.
Finding 5: there is a second silent surprise one section up
Enum casing is the one with the arresting wording, but the section immediately before it describes behaviour that is also invisible and also not an error.
Anthropic, verbatim: "required properties appear first, followed by optional properties"
Your schema's property order is preserved within each group, but required fields are hoisted above optional ones. The page's own example has a schema declaring notes, name, email, age with name and email required, and output ordered name, email, notes, age.
The advice is reasonable and ends with Anthropic, verbatim: "account for this reordering in your parsing logic."
For most consumers key order in a JSON object is irrelevant and this is a non-event. It stops being a non-event the moment something downstream is order-sensitive: a serialized form used as a cache key, a hash computed over the response body, a diff against a stored record, or a CSV writer taking its column order from the first object it sees. Same pattern as the enum: valid JSON, matching schema, not what you wrote.
Finding 6: this is not new, and I can bound that
I expected the enum caveat to be a recent addition, because a limitation that specific usually gets written down after somebody hits it. The dated snapshots say otherwise.
A third-party documentation diff monitor keeps dated captures of these pages. The structured outputs page resolves for 12 August, 1 September and 1 October 2026. The Enum value casing heading, the capitalization sentence, the no-special-stop_reason sentence and the case-insensitive advice are present on all three, and the only change I could find in the text around them between the earliest snapshot and today is a link being rewritten from a relative anchor to an absolute URL.
I read the diff lines rather than counting them, which mattered: the earliest capture shows the enum bullet twice, once as a removal and once as an addition, which looks like a change until you read both and see they are identical apart from the URL form.
So the caveat shipped with the feature, or near enough that I cannot distinguish it, and it has been stable for the two months I can observe. The strict tool use page has no capture at any date I tried, so I cannot say when its guarantee language was written. That is an absence of evidence and I am not treating it as evidence of anything.
One thing I will not claim: I have no idea how often this actually fires. The documentation says it may happen and gives no rate. A caveat that is two months old and barely discussed is consistent with "rare" and equally consistent with "common but silent", and silence is exactly what you would expect from a failure mode that produces a 200 and plausible output.
What I did not verify
Being explicit, because several of these would change how you act.
I have no API key and I measured no API behaviour. Everything above is a reading of published documentation, dated mirrors of that documentation, and third-party pages, as they stood on 10 October 2026. I did not send a single request. I did not observe an enum value come back with the wrong capitalization, and I cannot tell you how often it happens or on which models.
I did not test whether a Pydantic or Zod model actually raises on a case-drifted enum, which is the thing most teams would want to know first. It should, since the value is not in the declared set, but I am predicting that from the type systems and not reporting it.
I did not check whether the SDK helpers normalise casing on the way out. If they do, the exposure is narrower than I have described and limited to raw HTTP callers. The page says nothing either way and I did not read SDK source.
I did not verify the property-ordering behaviour against a real response, and I did not test whether it interacts with the enum drift.
I did not check the 17 models listed as supporting the feature for any per-model difference in this behaviour. The caveat is written without a model qualifier, so I am taking it as uniform, but "written without a qualifier" and "measured as uniform" are not the same thing.
And I did not cover the parts of this feature I have already written up. The 24-hour grammar cache, its invalidation rules, the zero-data-retention asterisk and the SDK stripping unsupported constraints are all in my earlier JSON schema retention piece, and that post's argument that the guarantee holds against a simplified schema is the closest thing I know to today's finding. They are not the same mechanism: there, an unsupported keyword is removed before compilation and your local validator catches the violation. Here, a fully supported keyword is compiled and the grammar drifts anyway, with no validator in the picture unless you kept one.
Credit where it is due on the loud paths: the Timo Labs guide already states that "a reply cut off at max_tokens, or a refusal, may not match the schema" and makes the broader point that strict mode "guarantees the shape, not the meaning". That is correct and well put, and it covers the two signalled failures. It carries nothing on the casing drift or the property reordering, which is the gap I went looking for.
If you want the typed Python path rather than the documentation archaeology, colleagues at AgentNotebook have a structured outputs tutorial that covers the code shape properly.
Postscript: I kept the validator. It is nine lines and it compares strings without caring about capital letters, which turns out to be the entire feature request.
Written by
M. PatelFrequently asked questions
Does Claude's structured outputs feature guarantee my output matches my schema?
It guarantees structural validity, and the documentation names four scenarios where the output may still not match your schema. Three of them give you a signal: a refusal sets stop_reason to refusal, a truncation sets stop_reason to max_tokens, and an unsupported JSON Schema feature returns a 400 error before generation. The fourth, enum value casing, returns a normal 200 response with no error and no special stop_reason.
What is the enum value casing caveat?
Anthropic documents that structured outputs do not guarantee the capitalization of string enum and const values, and that Claude may return a value differing from your schema only in capitalization, typically in the first letter of a word following a space. The response completes normally with no error. The documented mitigation is to compare enum values case-insensitively and to avoid enum values that differ only in capitalization.
Does the casing caveat apply to strict tool use as well as JSON outputs?
Yes. The caveat states in terms that it applies to both JSON outputs and strict tool use. It is published only on the structured outputs page. Measured on 10 October 2026, the strict tool use page contains zero occurrences of capitalization, casing or case-insensitive, links the json-schema-limitations anchor twice, and never links the invalid-outputs anchor where the caveat lives.
Which enum values are at risk?
The documented drift needs a string value containing a space, since it is described as affecting the first letter of a word following a space. Numeric enums are out of scope entirely, and single-word lowercase values such as celsius or fahrenheit are structurally immune. Multi-word display strings such as In Progress or Needs Review are the exposed shape, and they are also the most natural thing to write when an enum mirrors an existing status column.
Is this a recent change to the documentation?
No. Dated third-party captures of the structured outputs page from 12 August, 1 September and 1 October 2026 all carry the Enum value casing heading and its sentences, and the only change in the surrounding text across that window is a link rewritten from a relative anchor to an absolute URL. The caveat has been stable for the whole observable window.