Dani Reyes10 min read8 views

Priority Tier capacity is not extra capacity

An Anthropic Priority Tier commitment is drawn down alongside your ordinary rate limit, not on top of it. Reading the service tiers page next to the pricing page and the models list turns up a 20x burndown spread, a single cache read rate that only holds because of the exclusion list, and three of four current models excluded.

Flat schematic on a dark navy panel: one lime request node on the left, two connectors diverging into a lime priority capacity bar above and a slate standard rate limit bar below, both consumed to the same point and joined by an amber marker, with four rising blocks below for the burndown scale.
Flat schematic on a dark navy panel: one lime request node on the left, two connectors diverging into a lime priority capacity bar above and a slate standard rate limit bar below, both consumed to the same point and joined by an amber marker, with four rising blocks below for the burndown scale.
On this page

Quick answer

Checked on October 1, 2026. If you have an Anthropic Priority Tier commitment, the tokens per minute you committed to are not additional throughput. Anthropic's service tiers page states that a Priority request draws down your Priority capacity and your ordinary rate limit, and that the ordinary limit still declines the request if it would be exceeded. Priority Tier buys you position in the queue, not a bigger queue.

Three further things fall out of reading that page next to the pricing page and the models list, which I had never done in one sitting:

  • The capacity unit is a burndown unit, not a token, and the spread between the cheapest and dearest burndown is 20x.
  • The burndown table publishes a single, unqualified cache read rate. That rate is only correct because every model whose real cache read ratio differs is on the Priority Tier exclusion list.
  • Three of the four current models are on that exclusion list, including the one the docs tell you to start with.

The moment

I was not shopping for Priority Tier. I was finishing a note about rate limit headers nine days ago and I left a line in it saying I had not tested the anthropic-priority-input-tokens-* header family, because I have no Priority Tier access and was not going to pretend otherwise.

That line nagged. Not because I wanted to test the headers, but because while writing it I had quietly assumed something I never checked: that committed capacity is extra capacity. That if you buy 10,000 input tokens per minute of Priority, your ceiling goes up by 10,000.

It does not. And the sentence that says so is not on the rate limits page, which is where I would have looked.

Finding 1: a Priority request is charged to both meters

From Anthropic's service tiers page, as a note under the assignment rules:

Requests assigned Priority Tier pull from both the Priority Tier capacity and the regular rate limits. If servicing the request would exceed the rate limits, the request is declined.

Read that twice. Your Priority capacity is consumed, your standard rate limit is consumed, and the standard rate limit retains its veto. Committed capacity does not raise the ceiling. It decides who gets served first underneath the same ceiling.

That is a coherent product. Priority Tier is sold against overload errors, and the same page says the tier exists so that prioritisation helps minimise the server overloaded errors you would otherwise hit during peak times. Ordering is the thing being sold. My mistake was assuming headroom came with it.

What makes this easy to miss is where the sentence lives. I went looking on the rate limits page, which is the page you read to understand your ceiling. I counted the mentions: that page names Priority Tier twelve times, and all twelve are inside a single table of header definitions. The interaction is not stated there at all. It is one note on a different page, under a heading about how requests get assigned.

Finding 2: the unit on your contract is a burndown unit

A commitment is quoted in input tokens per minute and output tokens per minute. The service tiers page then explains that Anthropic counts usage against that capacity on a scale, not one for one. Quoting the list directly:

Cache reads as 0.1 tokens per token read from the cache

Cache writes as 1.25 tokens per token written to the cache with a 5 minute TTL

Cache writes as 2.00 tokens per token written to the cache with a 1 hour TTL

Plus 1.1 tokens per token for US-only inference on Claude 4.6 and later models, and one for one for everything else.

So the ratio between the cheapest and the most expensive input token is twenty to one. Ten thousand committed input tokens per minute is one hundred thousand tokens per minute if you are reading from cache, and five thousand if you are writing an hour-long cache entry.

The practical version: turning on a long cache is normally an unambiguous win, and against a Priority commitment it is two moves at once. The write side halves your effective committed throughput while it happens. The read side multiplies it by ten afterwards. Those are different minutes. If your traffic is bursty, the write burst and the read payoff do not land in the same rate limit window, and the burst is the part you bought Priority Tier to protect.

I wrote about schemas quietly becoming a second cached artifact a little over a week ago. This is the same shape of surprise on a different meter.

Finding 3: the single cache read rate survives because of the exclusion list

This is the one I did not expect, and it needed three pages open at once.

The burndown list publishes one cache read rate, 0.1, with no model qualification. The service tiers page justifies the whole table by saying:

These burndown rates reflect the relative pricing of each token type.

That is a falsifiable claim, so I went to check it. Anthropic's pricing page gives the cache read multiplier as 0.1x base input price, and then adds an exception in the same cell: 0.025x on Claude Fable 5.1 and Claude Mythos 5.1, and 0.05x on Claude Opus 5.5.

So there are three models where a cache read is not one tenth of base input price. On those three, an unqualified 0.1 burndown would misstate the price ratio by two to four times, and the table's stated rationale would be wrong.

It is not wrong, because all three of those models are on the Priority Tier exclusion list. The service tiers page ends with the line that Priority Tier is supported on all available Claude models except Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, and Claude Sonnet 5.

Set the two lists side by side:

Scroll to see more

ModelCache read ratioOn Priority Tier?
Claude Fable 5.10.025xNo, excluded
Claude Mythos 5.10.025xNo, excluded
Claude Opus 5.50.05xNo, excluded
Everything else0.1xYes

Every model that would break the single rate is a model you cannot use the tier on. The table is internally consistent, and it is consistent by exclusion rather than by qualification. I cannot tell you whether that is deliberate sequencing or coincidence, and I am not going to guess. What I can say is that it is load bearing: the moment a model with a 0.05x cache read becomes Priority eligible, that unqualified 0.1 stops reflecting relative pricing, and the note justifying the table stops being true.

Finding 4: three of the four current models are excluded

The exclusion list reads like a tail of older and specialist models until you line it up against the models overview. That page's comparison table carries exactly four current models. Three are on the exclusion list:

Scroll to see more

Current modelPriority Tier
Claude Fable 5.1Excluded
Claude Opus 5.5Excluded
Claude Sonnet 5.5Excluded
Claude Haiku 4.5Supported

The only current model that keeps Priority Tier is Haiku 4.5, which is the fastest model on the board and the one least likely to be sitting behind an overload error in the first place. Meanwhile the models page opens by telling you to start with Claude Opus 5.5 for most workloads, which is on the excluded list.

There is a second-order consequence here that I think matters more than the count. A commitment, per the same page, consists of input tokens per minute, output tokens per minute, a duration of 1, 3, 6 or 12 months, and a specific model version. A commitment is pinned to a model. If the newest models are not eligible, a multi-month commitment cannot be carried forward onto them. The upgrade path and the commitment point in opposite directions.

Finding 5: the default value selects a tier most orgs cannot buy

Anthropic The service tiers page opens with a warning that Priority Tier capacity commitments are no longer available for purchase, that existing holders can run to contract end, and that the page remains as a reference for them.

The request parameter did not go anywhere. service_tier still accepts auto and standard_only, and auto is still the documented default, meaning use Priority capacity if available and fall back otherwise.

For any organisation without a legacy commitment, that default resolves to the fallback every time. auto and standard_only are behaviourally identical for you, and the parameter is inert. It is not a bug and it costs nothing. It is just a live knob on every Messages request whose interesting branch is closed to new customers, and worth knowing before you spend an afternoon reasoning about which value to set.

If you are not sure which side of that line you are on, the page gives you a cheap oracle, and it keys on presence rather than value:

You can use the presence of these headers to detect if your request was eligible for Priority Tier, even if it was over the limit.

bash
curl -sS -D - -o /dev/null https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5-20251001","max_tokens":16,
       "messages":[{"role":"user","content":"ping"}]}' \
  | grep -i '^anthropic-priority-'

No lines back means no Priority eligibility on that key and model. There is a second, after the fact observable too: the response usage object carries a service_tier field telling you which tier actually served the request.

What I did not verify

  • I have no Priority Tier access. Every mechanism above is quoted from Anthropic's published documentation as of October 1, 2026, not observed against a live commitment. I did not send the header probe above against an eligible key, so I have not confirmed the headers appear when they should, only that the docs say presence is the signal.
  • I did not confirm the burndown rates empirically. That the 20x spread shows up as 20x in real capacity accounting is the documentation's claim, not my measurement.
  • I did not establish intent on Finding 3. The alignment between the cache read exceptions and the exclusion list is a fact about two published lists today. Whether one was built around the other, I do not know.
  • I did not check Bedrock, Microsoft Foundry, Google Cloud or Claude Platform on AWS. Everything here is the first party Claude API. Partner platforms set their own capacity terms and I have not read them.
  • I did not test whether service_tier: "standard_only" behaves differently from auto on a non eligible key. I am asserting they are behaviourally identical from the documented fallback, which is an inference and not a measurement.

One piece of dated provenance, since it is cheap to state and I used it: a public doc diff mirror carries this same page as of July 16, 2026. The purchase warning was already there in July. The only wording change I could see in the burndown list is a generalisation, from naming Claude Opus 4.6 and Claude Sonnet 4.6 explicitly to saying Claude 4.6 and later. The exclusion list, the 0.1 and the both meters note all predate my reading of them by at least two and a half months.

If you are handling the overload errors this tier exists to suppress rather than buying your way around them, the Python retry and backoff walkthrough on AgentNotebook covers the 429 and 529 side in detail. It does not touch service tiers, which is rather the point: that is the path the rest of us are on. And my earlier note on what the rate limit headers actually measure is the one that left this question open in the first place.

Postscript: I went in to close a loose end about six HTTP headers and came out having learned that the thing I assumed I was buying was never on sale.

D

Written by

Dani Reyes

Frequently asked questions

Does an Anthropic Priority Tier commitment raise my rate limit?

No. Anthropic's service tiers page states that requests assigned Priority Tier pull from both the Priority Tier capacity and the regular rate limits, and that the request is declined if servicing it would exceed the rate limits. The commitment buys priority ordering underneath the same ceiling rather than additional throughput. Checked October 1, 2026.

Why is a committed token not the same as a real token?

Priority Tier capacity is counted in burndown units. A cache read burns 0.1 tokens per token, a 5 minute TTL cache write burns 1.25, a 1 hour TTL cache write burns 2.00, and US-only inference burns 1.1 on Claude 4.6 and later models. The spread between the cheapest and dearest input token is therefore twenty to one, so a commitment quoted in tokens per minute has no fixed relationship to the tokens you can send.

Which Claude models still support Priority Tier?

As of October 1, 2026 the service tiers page excludes Claude Fable 5.1, Claude Mythos 5.1, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5 and Claude Sonnet 5. Set against the four current models in the models overview comparison table, that leaves Claude Haiku 4.5 as the only current model that keeps the tier.

Should I still set the service_tier parameter?

Priority Tier capacity commitments are no longer available for purchase, so for an organisation without an existing commitment the default value auto resolves to the standard fallback on every request and is behaviourally identical to standard_only. The parameter is inert rather than broken, and it costs nothing to leave at its default.

How can I tell whether my key is eligible for Priority Tier?

The service tiers page says you can use the presence of the anthropic-priority-input-tokens and anthropic-priority-output-tokens header families to detect whether a request was eligible, even when it was over the limit. Presence rather than value is the signal. The response usage object also carries a service_tier field reporting which tier actually served the request.

The container outlives its own expires_at

The Claude code execution tool returns an expires_at that is a short rolling value, not the real 30-day container limit. Three tool versions share two runtimes, two timeouts arrive in two different shapes, and web search quietly adds a second execution environment.

9 min read37

Exactly one tool call is no longer expressible

disable_parallel_tool_use means at most one tool call under tool_choice auto, and exactly one under any or tool. Claude Opus 5.5, Fable 5.1 and Mythos 5.1 reject any and tool with a 400, so the exactly-one contract is not expressible on them at all.

9 min read65

The Claude models endpoint has no retirement field

GET /v1/models returns a capability tree and no lifecycle field at all. Meanwhile two of the five platforms that serve Claude run a different retirement clock, and their model tables still list identifiers the Claude API retired months ago.

11 min read61