Claude's PDF page limit keys off a number you cannot see
The PDF support page states its page ceiling with a condition you cannot evaluate from that page. Three other statements of the same limit key it off the model instead. Measured on October 9, 2026, plus the arithmetic that makes 600 unreachable on a dense document.
On this page
Quick answer
On October 9, 2026, Anthropic's PDF documentation puts the per-request page ceiling at 600 (100 when the request's context window is under 1M tokens). Three other statements of the same limit, on two other pages, key it off the model instead: 100 per request on the API, for models with a 200k-token context window. Those are not the same condition. One you can look up; the other names a property of your request that nothing on the PDF page lets you evaluate.
It matters more than a wording nit, because on that page the 600 is the only number a reader takes away, and two other numbers on the same page make 600 unreachable for most real documents. If you are sending PDFs: your branch is decided by your model, not your request, and the practical ceiling is closer to 333 pages on a dense file.
The moment
A contractor sent me a 480-page scanned planning application and asked whether it would go through in one request. I had the page open. It said 600. I said yes.
Then I read the parenthetical properly. Six hundred pages, 100 when the request's context window is under 1M tokens. I went looking for which of those two numbers applied to us, and the page I was reading had nothing to tell me: it never names a model's context window, it never mentions a header, and the string 1M appears on it exactly once, inside that parenthetical.
So I went to the other pages. They answer the question, in different words, and the words matter.
Finding 1: one limit, four statements, two conditions, and the odd one out is the PDF page
The limit appears four times across three pages. I counted and read each one.
The PDF support page, in its requirements table:
600 (100 when the request's context window is under 1M tokens)
The context windows page:
A single request can include up to 600 images or PDF pages (100 for models with a 200k-token context window).
The vision page, twice. Once in a summary row, and once as a bullet list that is the clearest version of the rule anywhere in the documentation:
100 per request on the API, for models with a 200k-token context window.
600 per request on the API, for all other models.
Three of the four conditions are keyed to the model. One is keyed to the request. The model-keyed form is decidable: you chose a model, you can look its window up. The request-keyed form reads as a property of the payload you are about to send, which invites exactly the wrong inference, that a small request gets 100 pages and a big one gets 600.
I am not claiming the PDF page states a false limit. The numbers are the same on all four. The defect is that the one page scoped to the use case carries the one phrasing a reader on that page cannot act on, and the three actionable phrasings are on pages you have no reason to open when your question is about a PDF.
That page also closes the loop against itself. It says All active models support PDF processing. So every model is in scope, and the page sorts them by a property it never attributes to any of them. Measured on it: beta 0 occurrences, Haiku 0, Sonnet 0, price 0, premium 0, 100,000 0.
Finding 2: the branch is decidable, and the 100-page case is live
Off-page, it resolves cleanly. The context windows page lists fourteen models with a 1M-token window and adds Other Claude models, including Claude Sonnet 4.5 (deprecated), have a 200k-token context window.
Crossing that against the model deprecations status table, there are fifteen models in the Active state. Thirteen of them have a 1M window, so they get 600 pages. Exactly two Active models fall in the 100-page branch: claude-opus-4-5-20251101 and claude-haiku-4-5-20251001. One further model, claude-sonnet-4-5-20250929, is Deprecated with a retirement date of November 30, 2026, and is also in the 100-page branch until then.
So the narrow branch is real and reachable, not a leftover. Two of fifteen Active models is a small enough share that an engineer who assumes 600 will usually be right, which is precisely what makes being wrong about it expensive: it will work in every test you run on a current model and fail on the one service still pinned to Opus 4.5.
Finding 3: it is not a regression, and I cannot date it
The obvious story is that the request-keyed phrasing is a fossil. There used to be a beta header that opted a single request into a 1M window, and under that regime the request's context window was a caller-controlled property and the phrasing made sense.
The archived snapshots kill that story, at least inside the window I can see. On a dated snapshot of the PDF page from August 12, 2026, the cell already read 600 (100 when the request's context window is under 1M tokens), and the only change the diff recorded on that line was table padding. On the same date, the context windows page already read 100 for models with a 200k-token context window, with its only change a relative-to-absolute URL rewrite, and it already said For every model with a 1M-token context window, 1M is the default: you don't need a beta header. The vision page's bullets were already in their current form too.
So by the earliest date I can observe, the beta was already gone and both phrasings were already live and already divergent. They shipped divergent, and within my observation window they have never agreed. The origin predates what I can see, and I am not going to guess at it.
Two caveats on that evidence. Both snapshots come from one monitor, so they are mirrors of Anthropic's pages rather than independent records. And the snapshot tool renders a unified diff, so every changed line appears twice, once removed and once added. Counting occurrences on those pages tells you almost nothing; I read the lines.
Finding 4: 600 pages is only reachable in the bottom tenth of the page's own density range
The PDF page publishes the other number you need. Under text token costs, it says each page typically uses 1,500 to 3,000 tokens per page depending on content density.
Put that against a 1M-token window and the page ceiling stops being the binding constraint almost immediately.
A 1M window divided by 600 pages is 1,666.67 tokens per page. That is the heaviest average page you can send and still fit 600 of them, and it sits 11.1 percent of the way up the page's own stated 1,500-to-3,000 band. At the dense end of that band, 3,000 tokens per page, a 1M window holds 333 pages, which is 55.6 percent of the advertised ceiling. Six hundred pages at 3,000 tokens each is 1,800,000 tokens, or 1.8 times the entire window.
The page does flag this, and I want to be fair about it, because it flags it well. A tip directly under the requirements table says Dense PDFs (many small-font pages, complex tables, or heavy graphics) can fill the context window before reaching the page limit. That is the right warning in the right place.
My claim is narrower than contradicting it. The page states the tension qualitatively and never bounds it, and the bound is two numbers on that same page divided into each other. Both inputs are already published, forty lines apart. Nobody does the division, so the reader leaves with 600 and a vague sense that dense files are worse.
There is a second ceiling underneath both, which the page also states: the request itself is capped at 32 MB, and the API overview puts Bedrock at 20 MB and Google Cloud at 30 MB. For large files the page recommends uploading through the Files API and referencing by file_id, which keeps the payload small and does nothing at all about the token count. I have written about that header's response shape before and will not restate it here.
Finding 5: on one model, using the allowance crosses a 5x price cliff the PDF page never mentions
Claude Haiku 5.5 has a 1M-token context window, so it is in the 600-page branch. It is also the one model in the line-up priced by prompt length. From the pricing page:
Claude Haiku 5.5 is priced by prompt length: a request whose prompt is over 100,000 tokens pays higher prices.
The standard rate card has two rows for it, and every column multiplies by exactly five above the threshold. Base input goes from $0.10 to $0.50 per MTok. Five-minute cache writes go from $0.125 to $0.625. One-hour cache writes go from $0.20 to $1. Cache hits go from $0.01 to $0.05. Output goes from $0.50 to $2.50. The batch rows do the same thing at half the price, $0.05 to $0.25 on input.
Now bring the density band back. At 1,500 to 3,000 tokens per page, 100,000 tokens is somewhere between page 33 and page 67. That is 5.6 to 11.1 percent of the 600-page allowance the PDF page advertises. On the cheapest model Anthropic sells, you cross a five-fold price increase somewhere in the first tenth of the documented page budget, and the page that documents the budget says nothing about it.
The mitigation most people would reach for does not work either, and the pricing page is explicit about why: A request's prompt length counts all of its input tokens, including cache reads and cache writes. Caching the document does not keep you under the threshold, because the cached tokens still count toward the prompt length that selects your rate, and the cache-hit rate is itself one of the columns that went up fivefold. That is a different failure from the cache arithmetic I worked through on the batch side, where the question was which lifetime to buy; here caching is not the wrong choice, it just does not move the threshold.
The PDF page's only sentence about cost is Standard API pricing applies with no additional PDF fees. That is true in letter. There is no PDF surcharge. It is also the sentence a reader checks before sending a 400-page file, and on one model standard pricing is itself two rate cards with a cliff between them.
Proportion matters here and I do not want to oversell it. Nine hundred thousand input tokens at $0.50 per MTok is $0.45, against $0.09 if the same tokens billed at the lower rate, so the gap is about thirty-six cents on a single request. That is nothing on one call and it is $3,600 across ten thousand of them, which is roughly the volume Haiku 5.5 exists to serve. The structure is the finding, not the per-call number.
One dating note, offered narrowly. The Haiku 5.5 exception is absent from both the August 12 and September 1 snapshots of the context windows page and is present live, and the model's own row in the status table shows a retirement floor of October 7, 2027. That is consistent with a very recent addition, and I am not going to convert it into a release date I have not verified first-hand.
What I changed
I stopped reading the PDF page's parenthetical as a rule about requests and started treating it as a badly-worded rule about models. For the planning application, that meant checking the model first: on a 1M-window model the 480 pages are inside the page limit, and at scan density they are not inside the token limit, so it gets split regardless. We split it into four parts of 120 pages.
I also stopped using Haiku 5.5 for whole-document passes without checking the token estimate first. Not because it is the wrong model, but because the price I was reasoning about was the one above the line rather than the one on the rate card I had looked at.
What I did not verify
I have no Anthropic API key in this environment, so nothing here is a measurement of API behaviour. Every claim is a claim about what the documentation says, measured against the documentation on October 9, 2026, quoted rather than paraphrased, with the page it came from linked.
I did not send a 601-page PDF, or a 101-page PDF to a 200k-window model, so I have not established what the rejection looks like or which limit the error names. I also did not confirm that a request on a 1M-window model is actually accepted at 600 pages.
I did not establish that the PDF page's condition and the other three are functionally different in practice. My claim is about what a reader can evaluate from the page they are on, not that the two phrasings select different branches. If the platform does in fact compute a per-request window that can differ from the model's, then the PDF page is the precise one and the other three are the loose ones, and the whole finding inverts. Nothing I read says that is possible, and I could not rule it out from the documentation alone.
The 1,500-to-3,000 band is the vendor's own word typically, not a measured distribution, so every page-count figure I derived from it is an estimate carrying that word's uncertainty. I did not tokenize a real PDF to check it.
My two dated snapshots come from one monitor, they are mirrors rather than independent archives, and I did not corroborate either against a second source. I did not narrow the twenty-day gap between them, and the equivalent September 1 snapshots of the PDF and vision pages returned 404, which I am scoring as absence of evidence rather than evidence of anything.
I did not check whether any of this reproduces on the non-English documentation, and I did not check Bedrock, Vertex or Foundry beyond the request size limits quoted above. The PDF page has a separate Converse API section for Bedrock that I did not work through.
I deliberately did not compare this to how other vendors express the same limit. I have not read their current reference documentation, and writing that comparison from memory is how you end up citing a price that moved.
Adjacent ground I have already covered and am not restating: the behavioural side of the 1M window, where the question is not whether tokens fit but whether the model still uses them, is over here.
Postscript: the answer was on the page, divided by another number on the same page, forty lines down, and I needed a 480-page planning application to make me do the arithmetic.
Written by
Dani ReyesFrequently asked questions
How many PDF pages can I send to Claude in one request?
Six hundred on any model with a 1M-token context window, and 100 on a model with a 200k-token window. As of October 9, 2026 that means 13 of the 15 Active models get 600, while claude-opus-4-5-20251101 and claude-haiku-4-5-20251001 get 100. The PDF support page states the condition as a property of your request rather than of your model, which is harder to act on, but the vision and context windows pages both state the model-keyed version.
Is the 600-page limit actually reachable?
Rarely. Anthropic's PDF page says each page typically uses 1,500 to 3,000 tokens. A 1M-token window divided by 600 pages allows 1,666.67 tokens per page, which is near the bottom of that band. At 3,000 tokens per page a 1M window holds about 333 pages, roughly 56 percent of the advertised ceiling, so on a dense document the token limit binds long before the page limit does.
Does sending a large PDF change what I pay per token?
On Claude Haiku 5.5 it can. That model is priced by prompt length and every rate multiplies by five above 100,000 input tokens, including cache hits. At 1,500 to 3,000 tokens per page you cross 100,000 tokens somewhere between page 33 and page 67, which is inside the first tenth of the 600-page allowance. On the other models the 1M window is billed at standard pricing.
Does prompt caching keep a long PDF under the Haiku 5.5 threshold?
No. The pricing page states that a request's prompt length counts all of its input tokens, including cache reads and cache writes. Caching reduces what each token costs but not how many tokens count toward the threshold that selects your rate, and the cache-hit rate is itself one of the figures that goes up fivefold above the line.
Which models support PDF input?
The PDF support page states that all active models support PDF processing. That is why the page limit matters: every model is in scope, so the only thing deciding whether you get 600 pages or 100 is the context window of the model you picked.
Keep reading
The text editor tool prices a version the tool reference does not list (2026)
Claude's text editor tool publishes one token cost, keyed to text_editor_20250429, a version the canonical tool reference omits. The two versions it does list have no published figure, both follow-up links land on the wrong table, and the endpoint that would settle it is mentioned zero times on any tool page.
The Files API contradicts itself on expires_at, and the reference docs cannot warn you (2026)
Anthropic's Files API guide says expires_at appears on every file response. Its own migration table says the field is not returned under the legacy beta header. Dated mirrors put both sentences on the page since September 1, 2026, and neither API reference page can carry the warning.
Your Claude conversation has two prefixes, and only one of them errors
A thinking block stays valid against the bytes you sent. The prompt cache keys on the prompt the model renders. The rows where those two disagree cost money with no error.