Dani Reyes9 min read7 views

Exactly one tool call is no longer expressible

disable_parallel_tool_use means at most one tool call under tool_choice auto, and exactly one under any or tool. Claude Opus 5.5, Fable 5.1 and Mythos 5.1 reject any and tool with a 400, so the exactly-one contract is not expressible on them at all.

Flat schematic: four tool_choice mode bars on the left, the lower two marked blocked, feeding two outcome panels on the right of different sizes, one lit and one dimmed, with a modifier band at the far right.
Flat schematic: four tool_choice mode bars on the left, the lower two marked blocked, feeding two outcome panels on the right of different sizes, one lit and one dimmed, with a modifier band at the far right.
On this page

Quick answer

On 28 September 2026, disable_parallel_tool_use is not one setting. It is a modifier whose meaning is decided by the tool_choice type it sits inside. Under auto it buys you at most one tool call. Under any or tool it buys you exactly one. Those are different guarantees, and only one of them survives on the current flagship models, because any and tool now return a 400 on Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1. If your agent loop depends on getting exactly one tool call per turn, that contract is no longer expressible on those models at all.

Anthropic The two halves of this are each documented. What is not written down anywhere I could find is what happens when you put them together.

The moment

I was moving a small extraction loop from Claude Opus 5 to Claude Opus 5.5. The loop is boring by design: one turn, one tool call, parse the arguments, write a row. It had run on tool_choice: {"type": "any", "disable_parallel_tool_use": true} for months, which is the documented way to say "call exactly one tool, I do not care which".

The migration failed immediately with a 400. Fine, I thought, I will drop to auto and keep the flag. That request succeeded. The loop then ran for about forty minutes before it wrote a row with a null in it, because on one turn the model had answered in plain text and called nothing at all, and my parser had assumed a tool_use block would always be there.

That is not a bug in the flag. It is the flag doing exactly what it says under auto. I had swapped a guarantee for a weaker one without noticing, because the field name did not change.

Finding 1: the flag is not a top-level parameter, and its meaning is inherited

The first thing to get right is where the field lives. The parallel tool use page states it directly: "It is not a top-level request parameter. The effect depends on the tool_choice type."

So the shape is nested:

json
{
  "tool_choice": { "type": "auto", "disable_parallel_tool_use": true }
}

And the two documented effects are:

Scroll to see more

tool_choice typewith disable_parallel_tool_use: truecan the model call nothing?
auto (default)at most one tool callyes, it can answer in plain text
anyexactly one tool callno
toolexactly one tool callno
nonenot applicable, tools are blockedit must answer in text

The same page is explicit that under auto the model "can still answer in plain text without calling any tool". That sentence is the whole difference between the two rows, and it is the one I skipped.

This much is documented and other people have written it up accurately. I am flagging it as context, not as a discovery.

Finding 2: on the newest models, the "exactly one" row is unreachable

Here is the part I could not find stated anywhere as a single consequence.

"Exactly one tool call" requires tool_choice to be any or tool. Those two types are precisely what the current flagship models reject. The errors page gives the failure verbatim, a 400 invalid_request_error whose message reads "type tool and any are not supported for this model", and names Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1. It also notes the rejection applies on the token counting endpoint, so you cannot even price the request shape before sending it.

Put the two facts together and the result is a hole in the matrix:

Scroll to see more

modelat most oneexactly one
Claude Opus 5 and earlieravailableavailable
Claude Opus 5.5availablenot expressible
Claude Fable 5.1availablenot expressible
Claude Mythos 5.1availablenot expressible

The flag still takes true. The request still succeeds. You just silently get the weaker of its two meanings, because the stronger one was only ever reachable through a door that is now closed. Nothing in the response tells you which contract you ended up with. The only way to know is to read the tool_choice type you sent.

The practical consequence is that any code path written as "force a call, and only one" has to be rewritten as "ask for a call, handle the case where there is none". That is a parser change, not a parameter change, which is why swapping the type alone is not a migration.

Finding 3: the token table corroborates the gap structurally

The tool use overview carries a table of how many system prompt tokens the tool-use scaffolding costs. It has two columns, and the columns are split on exactly this boundary: one for auto and none, one for any and tool.

For Claude Opus 5 both cells are populated, at 286 tokens and 406 tokens. For Claude Opus 5.5 only the first is. The any and tool cell is empty.

I like this because it is not prose. The pricing table encodes the restriction as a missing value, which means the gating is not a late-added validation rule bolted on top. It reaches into how the request is assembled. On the models that do support forcing, it also tells you the price: forcing a tool costs about 120 extra system prompt tokens per request on Opus 5, on every call, whether or not the model would have picked the tool anyway.

Finding 4: the tidy explanation for why forcing went away does not survive the model lists

There is an attractive story available here and I want to record that I tested it and it failed.

The mechanism behind forced tool use is assistant prefill. The define tools page says that with any or tool "the API prefills the assistant message to force a tool to be used". Separately, the errors page documents that prefilling an assistant message is itself no longer supported on newer models, returning "This model does not support assistant message prefill. The conversation must end with a user message."

So the neat conclusion writes itself: prefill went away, and forced tool use went away with it, because forced tool use was built on prefill.

It does not hold. The two bans cover different sets. The prefill ban is scoped to Claude 4.6 and later, which is a large set. The forced tool use ban names exactly three models. Claude Opus 5 sits in the gap: it is inside the prefill ban and still accepts any and tool, which is why it has a populated 406 in the token table. If the one ban implied the other, Opus 5 could not exist in that state.

Whatever the internal mechanism is now, "it is prefill, therefore both died together" is not it. I am leaving this as a refuted hypothesis rather than replacing it with a better one, because I do not have evidence for a better one.

Finding 5: a server tool call can come back unrun, with no marker saying so

One more thing in the same area that bit me while I was reading around the parallel path, because it is also a case where the response shape does not announce what happened.

If you mix a server tool and one of your own client tools, the API can return with the server tool not executed. From the stop reasons page: "the API returns without running the server tool so that you can run the client tools first. There is no other marker for the state; detect it by checking each server_tool_use or mcp_tool_use block's id for a matching result block."

That last clause is the operative one. There is no flag. The detection is a join:

python
def unrun_server_tools(content):
    result_ids = {b.get("tool_use_id") for b in content if b.get("tool_use_id")}
    return [
        b["id"]
        for b in content
        if b.get("type") in ("server_tool_use", "mcp_tool_use")
        and b["id"] not in result_ids
    ]

Python If that list is non-empty, those calls have not happened yet. Return your client tool results and the server tool will run on the next turn.

What I did not verify

I did not run any of this against the API. Every claim above is read from the published documentation on 28 September 2026, not measured from live responses. In particular I have not confirmed by experiment that auto plus the flag really does permit a zero-tool turn on Opus 5.5 specifically, only that the documentation says auto permits it in general.

I have no access to Claude Mythos 5.1 or Claude Fable 5.1, so their inclusion in the gated list is taken from the errors page rather than reproduced.

The 286 and 406 token figures are the documented values for Claude Opus 5. I did not count tokens myself, and I did not check whether the empty cell for Opus 5.5 is a deliberate marker or a table that simply has nothing to put there.

I have not established why forced tool use was removed. Finding 4 rules out one explanation and offers no replacement.

Model availability and per-model restrictions change. Everything here is a snapshot of the docs on one day, and the model list in particular is the part most likely to be stale when you read this.

The interaction between tool_choice and prompt caching, and the way any and tool suppress the natural-language preamble, are both covered well in PromptAttic's tool_choice recipes; that is an adjacent surface to this post and I have deliberately not re-derived it here.

On this site, your JSON schema is a second cached artifact covers where strict goes and why it is not a field on tool_choice, and task budget caps a turn, not a task is the other post where a per-turn guarantee turned out not to be the per-task guarantee I had assumed.

Postscript: the field name is honest. It disables parallel tool use. It never promised to require serial tool use, and I read a promise into it that was only ever true in company.

D

Written by

Dani Reyes

Frequently asked questions

Is disable_parallel_tool_use a top-level parameter?

No. It is a field inside the tool_choice object, and the documentation states plainly that it is not a top-level request parameter. Its effect depends on the tool_choice type it is nested inside.

What is the difference between at most one and exactly one tool call?

With tool_choice auto plus disable_parallel_tool_use true, Claude calls at most one tool and may still answer in plain text with no tool call at all. With tool_choice any or tool plus the same flag, Claude calls exactly one tool. Only the second is a guarantee that a tool_use block will be present.

Why does tool_choice any return a 400 on Claude Opus 5.5?

Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1 do not support forced tool use. Sending tool_choice of type any or tool to any of them returns a 400 invalid_request_error, including on the token counting endpoint. Only auto and none are accepted.

How do I get exactly one tool call on Claude Opus 5.5?

You cannot express that contract directly. The exactly-one guarantee requires tool_choice any or tool, and those types are rejected. You can set auto plus disable_parallel_tool_use true to cap the response at one call, but you must then handle turns where no tool is called at all.

How do I tell whether a server tool actually ran?

Check whether each server_tool_use or mcp_tool_use block has a matching result block with the same id. The documentation states there is no other marker for the state. If a block has no matching result, that server tool was not run and the API returned early so you could execute your client tools first.

The Claude models endpoint has no retirement field

GET /v1/models returns a capability tree and no lifecycle field at all. Meanwhile two of the five platforms that serve Claude run a different retirement clock, and their model tables still list identifiers the Claude API retired months ago.

11 min read26