M. Patel10 min read7 views

Claude's token counter stops at the advisor's first call

The advisor tool is the only server tool Anthropic's count_tokens endpoint accepts, and the count covers the executor's first sampling call only. The top-level usage object is executor-only too, so both default meters miss the advisor sub-inference. Arithmetic on Anthropic's own published example puts 75.2 percent of output tokens outside the headline figures.

Flat schematic of one Claude Messages request: three iteration blocks on a timeline, two small executor blocks and one large advisor block, with a short bracket covering only the first executor block and a second bracket covering both executor blocks but skipping the advisor entirely.
Flat schematic of one Claude Messages request: three iteration blocks on a timeline, two small executor blocks and one large advisor block, with a short bracket covering only the first executor block and a second bracket covering both executor blocks but skipping the advisor entirely.
On this page

Quick answer

Checked on October 2, 2026. The advisor tool is the only server tool that Anthropic's Anthropic token counting endpoint will accept. Every other server tool makes the request fail outright. That reads like a courtesy until you read the scope line attached to it: the count covers the executor's first sampling call, and nothing else.

The advisor sub-inference is not in that number. It runs on a second, usually more expensive model, and it is fed your entire transcript. It is also not in the top-level usage object you would reach for after the request returns, because those fields are executor-only by design.

So both of the obvious places to look are measuring the same half of the request, in the same direction, and the two documentation pages that each explain one half never point at each other.

The moment

I was costing out a migration to the advisor pattern on a long-horizon coding loop. The pitch is good: run a cheap executor, let it consult an expensive advisor at the two or three moments where the plan actually matters, and pay the premium rate only for the advice.

I did what I thought was the careful version. I counted the request before sending it, using the endpoint that exists for exactly this. I got a number. I ran the real thing. I read usage.input_tokens and usage.output_tokens off the response and compared them to my estimate.

They broadly agreed. I wrote in my notes that the advisor was cheaper than expected and moved on.

They agreed because I had built two meters that both measure the executor. The estimate excluded the advisor and the receipt excluded the advisor, so of course they matched. Two instruments agreeing is only evidence when they are independent, and these two share a blind spot.

Finding 1: The one countable server tool is the one the count describes worst

The token counting endpoint rejects server tools. It also rejects the MCP connector, and image or document blocks that use a url or file source. The advisor tool is carved out as the single exception, and the carve-out ships with a qualifier.

Anthropic, verbatim: "Token counting supports client tools and the advisor tool. Requests that include other server tools return an error. For the advisor tool, the count covers the executor's first sampling call only."

That is from the token counting page.

Sit with the shape of that. Every other server tool fails loudly. You send the request, you get an invalid_request_error, and you are left in no doubt that this endpoint cannot help you. The advisor tool succeeds quietly and hands back a number that describes a fraction of the work.

A rejection is honest. A partial count is not, unless you read the sentence that scopes it.

And the scoping is tighter than it first looks. It is not "the count excludes the advisor". It is the executor's first sampling call. An advisor turn is a sequence of executor iterations with advisor sub-inferences between them. Everything after the first executor iteration is outside the number too.

Finding 2: The documented fallback is blind in the same direction

The token counting page anticipates that it cannot count some requests, and it tells you what to do instead.

Anthropic, verbatim: "For requests that use server tools or MCP servers, the Messages API response reports the tokens used in its usage object."

That is the remedy, on the same page. Count what you can; for the rest, send the request and read the receipt.

Now the advisor tool page, on the subject of that receipt.

Anthropic, verbatim: "Top-level usage fields reflect executor tokens only. Advisor tokens are not rolled into the top-level totals because they are billed at a different rate."

Both sentences are correct. Together they leave a hole. The page that sends you to the usage object does not mention that for the advisor tool the usage object omits the advisor, and the page that documents the omission does not mention that another page just recommended the field as the fallback.

The information you need is a one-line array lookup away. It is usage.iterations, and I will come back to why I of all people should have known that. But nothing in the pre-flight path tells you to go there, and the field a cost meter reaches for by default is the one that is executor-only.

Finding 3: Anthropic's own example sizes the gap, and the arithmetic checks out

The advisor page publishes a worked usage object. I did the sums on it rather than taking the shape on trust.

json
{
  "usage": {
    "input_tokens": 1760,
    "output_tokens": 531,
    "iterations": [
      { "type": "message",          "input_tokens":  412, "output_tokens":   89 },
      { "type": "advisor_message",  "input_tokens":  823, "output_tokens": 1612,
        "model": "claude-opus-5" },
      { "type": "message",          "input_tokens": 1348, "output_tokens":  442 }
    ]
  }
}

The two message iterations are the executor. 412 plus 1348 is 1760, which is the top-level input_tokens exactly. 89 plus 442 is 531, which is the top-level output_tokens exactly. The top-level figures are the executor iterations summed, to the token, and nothing else.

The advisor_message iteration contributes 823 input and 1612 output, and appears in neither total.

Run that out. Total output across the whole request is 2143 tokens, of which 1612 are invisible at the top level. That is 75.2 percent of the output tokens, and they are the ones billed at the advisor model's rate rather than the executor's. The advisor's output on its own is 3.04 times the entire top-level output figure.

On the input side the distortion is milder but still real: 823 of 2583 input tokens, a little under a third, sit outside the headline.

I want to be precise about what this is and is not. It is one illustrative example in a documentation page, not a benchmark, and the real ratio on your workload depends entirely on how often the executor consults the advisor. What the example does establish, because the arithmetic is exact rather than approximate, is that the top-level fields are constructed to exclude the advisor. This is not drift or rounding. It is the documented design, and the page says so in the sentence quoted above.

Finding 4: I had already written this rule down, under the wrong heading

Three weeks ago I published a post about Claude's refusal fallback behaviour, and one of its findings was that the per-attempt record is usage.iterations, not the top-level usage. When a model declines and another serves the turn, the top-level counts describe only the attempt that produced the returned message, and tokens from different models are never summed into one field.

I filed that under exception handling. It felt like a fallback concern, something that bites when the unusual path fires, and I fixed my cost meter for that case and considered the matter closed.

It is not an exception-path rule. It is the rule for any turn where more than one model did work, and the advisor tool is that situation by design, on the success path, on every turn that uses it. Refusal fallback is the rare version. The advisor is the version you are opting into deliberately and will see constantly.

Generalising from the case where you found something is the cheap part. Noticing which other cases it covers is the part I skipped, and the cost of skipping it was that I built the correct fix into one code path and left the default meter wrong everywhere else.

The honest summary: my fallback post got the mechanism right and the scope wrong.

Finding 5: What the counter could not tell you even if it tried

It would be easy to read all of the above as an oversight that a future endpoint revision will close. I do not think most of it is, and the reason is worth stating, because it changes what you should build.

The advisor's input is not a parameter you supply. It is assembled server-side from the conversation as it stands at the moment the executor decides to ask.

Anthropic, verbatim: "That transcript includes your system prompt, the tool definitions, the prior turns and tool results, and the text the executor has produced so far in this turn."

That is on the advisor tool page, and the same page adds that "The advisor itself runs without tools and without context management."

Put those together. The advisor's prompt is the whole conversation including text that does not exist yet when you count, and the mechanism that would trim a growing transcript is explicitly switched off for it. The max_uses parameter caps advisor calls per request, but it is a ceiling rather than a prediction, and whether the executor reaches for the advisor once or never is a decision made mid-generation.

So a pre-flight count of an advisor request is not a slightly incomplete estimate. The quantity it would need to estimate is not determined at the time it runs. Reporting the first executor sampling call is arguably the only honest thing the endpoint can do.

The practical consequence is that for this tool the receipt is the only instrument, and it has to be read per iteration. If you are tracking spend, read usage.iterations and branch on type, because advisor_message entries carry their own model field and bill at that model's rate.

What was already written down

Worth saying plainly, because I checked before writing rather than after. The mechanics of the advisor tool are well covered already. getclaudeskills has a thorough walkthrough of the parameters, the result variants, the caching layers, and the fact that top-level max_tokens does not bound the advisor, and it reports the same usage.iterations structure I quote above. A separate certification-prep write-up covers the token counting endpoint's rejection list and the tokenizer change almost sentence for sentence with the docs.

What I could not find anywhere, and what this post is actually about, is the join: the pages covering the advisor tool do not mention the token counting endpoint, and the pages covering token counting do not mention the advisor tool beyond listing it as the exception. Each half is documented. The consequence of holding both at once does not appear to be.

On the token counting endpoint itself, my colleague over at PromptAttic has five recipes for using it well, which is the adjacent and more generally useful piece. This post is the footnote to it that says: one tool on that endpoint returns a number that means something narrower than you think.

What I did not verify

  • I did not run an advisor request. Everything here is read from the live documentation on October 2, 2026, plus arithmetic on Anthropic's own published example. I have not confirmed empirically that a pre-flight count of an advisor request equals the first message iteration exactly, only that the documentation says the count covers that call.
  • I did not test the interaction with max_uses. Whether setting it changes the returned count at all is unknown to me, and I would guess not, but a guess is not a measurement.
  • I did not check Bedrock or Vertex. The advisor tool is a beta on the Claude API and partner platforms routinely diverge on beta features.
  • I did not measure the real-world ratio. The 75.2 percent figure is arithmetic on one documentation example, not a finding about typical workloads.
  • I did not read the streaming path closely. The docs say the advisor result arrives in a single content_block_start event with a following message_delta carrying the updated iterations, which implies the per-iteration data is available to streaming clients too, but I have not exercised it.

Postscript: the thing that actually cost me a week was not the missing tokens, it was two meters agreeing. I now treat agreement between instruments as a prompt to check whether they are independent, rather than as confirmation.

M

Written by

M. Patel

Frequently asked questions

Does the Claude token counting endpoint support the advisor tool?

Yes, and it is the only server tool it supports. Every other server tool, plus the MCP connector and image or document blocks using a url or file source, returns an invalid_request_error. But the documentation scopes the advisor count to the executor's first sampling call only, so the advisor sub-inference and every later executor iteration are excluded from the number.

Why do advisor tokens not appear in the top-level usage object?

Because they are billed at a different model's rate. The advisor tool documentation states that top-level usage fields reflect executor tokens only and that advisor tokens are not rolled into the totals. They are reported separately in the usage.iterations array, where entries with type advisor_message carry their own model field.

Where do I read what an advisor call actually cost?

The usage.iterations array on the Messages API response. Branch on the type field: message entries are executor iterations billed at the executor model's rate, and advisor_message entries are the advisor sub-inference billed at the advisor model's rate. A cost meter that reads only usage.input_tokens and usage.output_tokens will miss the advisor entirely.

Can I estimate an advisor request's cost before sending it?

Not fully, and probably not in principle. The advisor is fed the executor's transcript as it stands when the executor decides to consult it, including text the executor has produced so far in that turn, so the input does not exist at count time. Whether the executor calls the advisor at all is also decided mid-generation, and max_uses is a per-request ceiling rather than a prediction.

Is this the same issue as the refusal fallback usage reporting?

It is the same structure on a different path. Refusal fallback also reports per attempt in usage.iterations because more than one model did work. The difference is that fallback is an exception path, while the advisor tool puts two models on the success path deliberately, on every turn that uses it.

Priority Tier capacity is not extra capacity

An Anthropic Priority Tier commitment is drawn down alongside your ordinary rate limit, not on top of it. Reading the service tiers page next to the pricing page and the models list turns up a 20x burndown spread, a single cache read rate that only holds because of the exclusion list, and three of four current models excluded.

10 min read37

The container outlives its own expires_at

The Claude code execution tool returns an expires_at that is a short rolling value, not the real 30-day container limit. Three tool versions share two runtimes, two timeouts arrive in two different shapes, and web search quietly adds a second execution environment.

9 min read59

Exactly one tool call is no longer expressible

disable_parallel_tool_use means at most one tool call under tool_choice auto, and exactly one under any or tool. Claude Opus 5.5, Fable 5.1 and Mythos 5.1 reject any and tool with a 400, so the exactly-one contract is not expressible on them at all.

9 min read75