The container outlives its own expires_at
The Claude code execution tool returns an expires_at that is a short rolling value, not the real 30-day container limit. Three tool versions share two runtimes, two timeouts arrive in two different shapes, and web search quietly adds a second execution environment.
On this page
Quick answer (2026): The Claude code execution tool hands back a container object with an expires_at timestamp, and that timestamp is not when your container expires. The real limit is 30 days from creation; expires_at is a shorter rolling value. A container goes quiet after roughly five minutes of inactivity, gets checkpointed, and restores when you send its ID again inside the 30-day window. Read from the code execution tool docs on 29 September 2026.
I lost a container I did not need to lose.
The moment
I had a long-running analysis job split across requests. First request builds a dataframe, writes it to /tmp, hands back a container ID. Later requests reuse that ID and keep working on the same files. This is the documented pattern and it works well.
I built a small guard around it. Before reusing a container I checked expires_at, and if the timestamp was close I threw the container away and started clean. That felt like the responsible thing to do. Nobody wants a request to fail on a dead sandbox in the middle of a job.
My guard fired constantly. It kept telling me containers were about to die, and it was wrong every time.
The docs say it plainly, in a sentence I had read and not absorbed.
Verbatim, from the code execution tool page: "The expires_at timestamp in the response's container object is a shorter rolling value and doesn't report the 30-day limit."
I had been rebuilding state every few minutes to protect against an expiry that was a month away.
Finding 1: expires_at is a rolling value, not a deadline
There are two clocks and only one of them is in the response.
The one you can see is expires_at. It rolls. It is short. It moves as you keep using the container, and it tells you very little about how long your files will survive.
The one that matters is not in the object at all: containers expire 30 days after creation. Between those two numbers sits a checkpoint. After about five minutes of inactivity the container is checkpointed, and sending a request with its ID inside the 30-day window restores it.
So the lifecycle has three states, and the response surfaces a timestamp that describes none of them cleanly:
created ──▶ active ──▶ checkpointed (~5 min idle) ──▶ restorable (up to 30 days) ──▶ expired
expires_at rolls somewhere in here
The practical consequence is the one that cost me: a checkpointed container is not a lost container. It is the normal resting state of any workflow with a human in the loop, or a queue, or a nightly schedule. Treating a short expires_at as a reason to start over throws away exactly the state the reuse feature exists to preserve.
Finding 2: an expired container fails loudly, which is what you want
The other half of my guard was defensive for a different reason. I assumed that passing a stale container ID might silently give me a fresh, empty container, and that my code would then happily read a file that was not there and produce confident nonsense.
That is not what happens. An expired container cannot be reused, and a request that references one returns an error rather than restoring it. The documented recovery is to send the request again without the container parameter, which gets you a new one.
That is a much better contract than the one I was defending against. It means the correct shape is a retry, not a pre-check:
def run(client, payload, container_id=None):
kwargs = dict(payload)
if container_id:
kwargs["container"] = container_id
try:
return client.messages.create(**kwargs)
except anthropic.APIStatusError:
# stale container: drop it and take a fresh one
return client.messages.create(**payload)
Reading
expires_at to decide whether to reuse is guessing. Trying the reuse and handling the failure is knowing. The failure is cheap and it is explicit.
Finding 3: three tool versions, two runtimes
The tool has three current versions and every supported model accepts all three. I expected three behaviours. There are two.
Scroll to see more
| Tool type | Runtime | What it adds |
|---|---|---|
code_execution_20250825 | A | Bash commands and file operations |
code_execution_20260120 | B | REPL state persistence, programmatic tool calling |
code_execution_20260521 | B | Nothing. Same runtime as 20260120 |
code_execution_20260521 is the same runtime as code_execution_20260120. The difference is in the tool description: the newer version tells Claude about the 90-second wall-clock limit on each Python cell, so Claude can budget long-running cells.
That is a version bump whose entire payload is a prompt change. No new capability, no new response shape, nothing to migrate. What it buys you is a model that knows about a constraint it was previously discovering by hitting it.
I find this genuinely interesting as an API design decision, because it is the vendor treating the tool description as a versioned artifact in its own right. The schema is identical. The instructions are not.
Finding 4: two timeouts, and only one of them is an error
This is the finding I would have debugged for an hour.
There are two separate execution limits and they surface in two completely different shapes.
A whole tool invocation that runs past the maximum execution time returns an execution_time_exceeded error code, in an error block, alongside siblings like unavailable, invalid_tool_input and too_many_requests. That is the shape you would expect.
A single Python cell under programmatic tool calling has its own 90-second wall-clock limit, and it does not do that at all.
Verbatim, from the same tool versions section: "A cell that exceeds the limit returns a normal code execution result with a non-zero return_code and a detection_timeout status message in its output."
A normal result. Not an error block. If you are branching on error codes to decide whether something timed out, a cell timeout walks straight past you wearing the costume of a command that simply exited non-zero.
{
"type": "bash_code_execution_tool_result",
"content": {
"return_code": 1,
"stdout": "... detection_timeout ..."
}
}
Two limits, two shapes, one word in common. The 90-second limit is per cell; execution_time_exceeded is per invocation. They are not the same event and they do not arrive the same way.
Finding 5: accepting a tool type is not supporting it
Claude Haiku 4.5 accepts code_execution_20260120 and code_execution_20260521. Send either and you get no complaint, no warning, no 400.
You also do not get what you asked for. Programmatic tool calling and the REPL state persistence that depends on it are not available there, so on Haiku 4.5 the newer versions behave like code_execution_20250825.
This is a shape I keep running into and it is worth naming: acceptance and support are different guarantees, and the API only enforces the first. A request that validates tells you the field was well formed. It does not tell you the capability exists.
It also interacts badly with the version table above. If you route cheap work to Haiku and expensive work to a larger model, and both branches send code_execution_20260521 because that is the newest, then one branch has REPL persistence and the other silently does not, and the tool type in your logs is identical on both.
Finding 6: web search hands you a second execution environment
The last one surprised me most because it arrives without being asked for.
The current web search and web fetch tools require code_execution_20260120 or later as their code execution version. Add web search to a request and code execution comes with it, enabled automatically.
If your application already provides its own client-side shell tool, you now have two execution environments in one conversation: Anthropic's sandboxed container, and whatever you run locally. The docs call this a multicomputer environment and are candid about the failure mode.
Verbatim, from the same page: "Claude can sometimes confuse these environments, attempting to use the wrong tool or assuming state is shared between them."
Variables, files and state do not persist between them. A file written in the container is not on your disk, and the container has no internet access at all, so it cannot go and fetch it either. If you are running your own sandbox for agent file tools, that boundary is now something you have to state explicitly rather than assume, and it is worth being deliberate about how you scope a local agent sandbox when a second one appears next to it.
There is a protocol detail underneath this that is easy to misread. When Claude calls one of your client tools alongside code execution, the API returns the code execution call without its result, and the result arrives in a later response after you send back your
tool_result blocks.
That is the same behaviour I found from the other direction last week: a server tool call with no matching result block has not failed, it has not run yet. Two different documentation pages now describe that state, and neither gives it a flag. You detect it by joining IDs, and the stop reasons page says so explicitly.
What I did not verify
- I did not run any of this against the live API. Every claim here is read from the vendor documentation on 29 September 2026, not measured. In particular I did not observe a real
detection_timeoutpayload, so the exact field placement in my JSON sketch above is my reading of the prose and not a captured response. - I did not test container restoration near the 30-day boundary. I have no way to age a container a month inside one session, so the checkpoint-and-restore behaviour is documented behaviour I am reporting, not behaviour I have seen.
- I did not cover billing, deliberately. The pricing rules for this tool are unusual and worth knowing, but they are already covered accurately elsewhere and none of the findings above depends on them.
- I did not benchmark whether the
code_execution_20260521description change actually alters how Claude budgets cells. The documented difference is a description, and whether that measurably changes behaviour is an empirical question I did not answer. - I did not examine the memory, bash or text editor tools as client-side tools in their own right, only the file operations that the code execution tool exposes. Those are separate surfaces with separate error codes.
Postscript: the guard I wrote to protect a container was the only thing that ever destroyed one.
Written by
M. PatelFrequently asked questions
Does expires_at tell me when my code execution container dies?
No. The documentation states that expires_at is a shorter rolling value that does not report the 30-day limit. Containers expire 30 days after creation. After roughly five minutes of inactivity a container is checkpointed, and sending a request with its ID inside the 30-day window restores it, so a short expires_at is not a reason to discard a container.
What happens if I reuse a container ID that has expired?
The request returns an error rather than silently giving you a fresh container. The documented recovery is to send the request again without the container parameter, which gets you a new one. This means the correct pattern is to attempt the reuse and handle the failure, rather than pre-checking a timestamp.
What is the difference between code_execution_20260120 and code_execution_20260521?
They are the same runtime. The only difference is the tool description: the 20260521 version tells Claude about the 90-second wall-clock limit on each Python cell in programmatic tool calling, so Claude can budget long-running cells. There is no new capability and no new response shape.
How does a 90-second cell timeout appear in the response?
Not as an error. A cell that exceeds the limit returns a normal code execution result with a non-zero return_code and a detection_timeout status message in its output. That is distinct from execution_time_exceeded, which is an error code returned when a whole tool invocation exceeds the maximum execution time.
Does Claude Haiku 4.5 support the newer code execution tool versions?
It accepts the 20260120 and 20260521 tool types without error, but programmatic tool calling and the REPL state persistence that depends on it are not available there, so the newer versions behave like code_execution_20250825. Accepting a tool type and supporting its capabilities are different guarantees.
Keep reading
Exactly one tool call is no longer expressible
disable_parallel_tool_use means at most one tool call under tool_choice auto, and exactly one under any or tool. Claude Opus 5.5, Fable 5.1 and Mythos 5.1 reject any and tool with a 400, so the exactly-one contract is not expressible on them at all.
The Claude models endpoint has no retirement field
GET /v1/models returns a capability tree and no lifecycle field at all. Meanwhile two of the five platforms that serve Claude run a different retirement clock, and their model tables still list identifiers the Claude API retired months ago.
Claude refusal fallback is sticky, and the second turn looks like nothing happened
A Claude refusal is an HTTP 200 with a fallback subsystem behind it. The handoff is marked by a content block, and after the first fallback the routing goes sticky and that marker disappears while the fallback model keeps answering.