Code Mode is a way for an AI agent to use its tools by writing code. The agent doesn't emit one tool call, wait for the result, and then emit the next. It writes a short script that calls the tools as functions, runs that script in a sandbox, and only the final output goes back into the model's context. Cloudflare coined the name "Code Mode" for MCP tools in 2025. Anthropic and OpenAI now ship the same idea in their APIs as "programmatic tool calling", and coding agents such as Codex and Pi have their own versions.
This guide covers what Code Mode for MCP actually changes, which implementations exist as of October 2026, when it saves tokens and when it doesn't, and the security trade-off you take on by running model-written code.
How Code Mode differs from normal tool calling
In classic tool calling, every step goes through the model. The model asks for tool A, your app runs it, and the full result goes back into the context. Then the model reads that result and asks for tool B, often copying data from A's output into B's input. A ten-step job means ten model turns, and every intermediate result takes up context space.
Code Mode collapses those steps. The agent gets one tool that runs code, plus a typed API that describes the other tools. The model writes a program that can loop, branch, run calls in parallel, filter big responses, and pass data from one call to the next without reading it. The sandbox runs the program and returns only what the script prints or returns.
Two arguments drive the approach:
- Models write code better than they emit tool calls. Cloudflare's original Code Mode post argues that models have seen huge amounts of real TypeScript but only a small, synthetic set of tool-call examples. So they handle more tools, and more complex ones, when those tools look like an ordinary API.
- Intermediate results stay out of the context. Anthropic's engineering post Code execution with MCP shows a meeting transcript copied from Google Drive into Salesforce. With direct calls, the whole transcript passes through the model twice. With code, it moves between the two tools inside the sandbox.
MCP still matters here. It gives the harness a standard way to connect to a server, list its tools and handle authorization. Code Mode just changes how the model is shown those tools. If MCP itself is new to you, start with our MCP server tutorial and the MCP vs A2A explainer.
Where Code Mode is available in 2026
| Implementation | Language the model writes | Where the code runs | Status |
|---|---|---|---|
Cloudflare @cloudflare/codemode | TypeScript/JavaScript | Dynamic Workers (V8 isolates) or a browser iframe | npm 0.5.3, MIT |
| Claude API programmatic tool calling | Python | Anthropic code execution container | Needs code_execution_20260120 or later |
| OpenAI Responses API programmatic tool calling | JavaScript | Fresh, isolated V8 runtime hosted by OpenAI | On by default in the Agents API |
Codex CLI features.code_mode | JavaScript | Codex's V8 runtime | Under development, off by default |
| Pi 1.0 Codemode | JavaScript | QuickJS inside a WASM runtime | On when MCP is enabled |
Cloudflare Code Mode
Cloudflare's package turns tool schemas into TypeScript declarations and gives the model a single code tool. Per the Code Mode API reference, it has entry points for the AI SDK (the Vercel AI SDK), MCP, TanStack AI, the browser and Vite. codeMcpServer() wraps an existing MCP server so it exposes one code tool. openApiMcpServer() builds an MCP server with just search and execute tools over an OpenAPI spec, and the host keeps the credentials outside the sandbox.
The default executor runs each script in a Dynamic Worker with a 60-second timeout, and outbound network access is blocked unless you pass a proxy. The newer runtime API adds durable executions that can pause for human approval on chosen methods, roll back actions that define a revert function, and save working scripts as reusable snippets. Dynamic Workers are only available on the Workers Paid plan.
Claude: programmatic tool calling
In the Claude API you add the code execution tool and set allowed_callers on each tool that code may call. Claude writes Python, and your tools show up as async functions. When the script calls one, the API pauses and returns a tool_use block. You send back the result, and the script carries on. Per Anthropic's docs, those intermediate results don't count toward your input or output tokens. Only the final output does.
Some limits to know. Tools from Anthropic's MCP connector can't be called this way, and neither can tools with strict: true. A pending call times out after about four minutes, and idle containers are reclaimed after about five. Anthropic also warns that allowed_callers guides the model but isn't a security boundary. Pricing follows code execution pricing. This is the API feature. Claude Code is a separate product.
OpenAI: programmatic tool calling
In the Responses API you add the programmatic_tool_calling hosted tool and mark eligible tools with allowed_callers: ["programmatic"] (or allow both direct and programmatic calls). According to OpenAI's guide, function, custom, MCP, apply_patch, shell and code interpreter tools can all be called from a program. Each program runs in a fresh V8 runtime with no Node.js, no direct network access, no general file system and no state kept between runs. An MCP tool's approval policy can pause the program until you approve the call. The Agents API turns programmatic tool calling on by default.
Codex and Pi
The Codex configuration reference lists features.code_mode.enabled and describes it as under development and off by default. Two related settings, excluded_tool_namespaces and direct_only_tool_namespaces, keep chosen tools out of code mode. To experiment with it in Codex:
[features.code_mode]
enabled = true
Pi, the minimal open-source coding harness, added MCP support through Codemode in its 1.0 release. Armin Ronacher, who works on Pi, explains in What is Codemode that it runs in QuickJS inside WASM, with no network, file system or timers and limited memory. It is on by default only when MCP is enabled.
When Code Mode saves tokens, and when it doesn't
The savings depend on the shape of the job. Anthropic published its own internal numbers for programmatic tool calling on a production Claude model:
- On a 75-tool project-management benchmark, billed input tokens fell by roughly 38% with no change in accuracy.
- On τ²-bench, where each turn makes only one or two sequential calls, scores were unchanged but cost was about 8% higher.
- Across production traffic, requests with 10 to 49 tools typically saved 20% to 40% of tokens.
So Code Mode is a strong fit for:
- Fan-out: checking 50 endpoints, or looking up 20 records, in one script.
- Big results you can shrink: filter a 10,000-row sheet down to the five rows that matter before the model sees it.
- Predictable pipelines: call A, use a field from it to call B, then aggregate.
- Many tools: load a tool's description only when you need it. Anthropic's post suggests a file tree of tool files or a
search_toolstool.
And a weak fit for:
- One or two small calls, where starting a sandbox and writing a script costs more than it saves.
- Steps where every result needs fresh judgment from the model before the next call.
- Writes and approval-sensitive actions. OpenAI's guide recommends direct tool calls by default here, so there's a clear authorization point.
OpenAI's advice is to measure: run direct calls as a baseline, then compare accuracy, tokens, latency and tool calls on real tasks.
Make your MCP server Code-Mode friendly
If you maintain an MCP server, a few changes help agents script against it:
- Declare an
outputSchemaand returnstructuredContent. The MCP tools spec supports both. A script needs predictable JSON, not prose. OpenAI and Anthropic both ask for documented output formats for the same reason. - Keep results consistent. Ronacher describes an agent that tests a call on five items and then fails at full batch size because the server changed its response shape.
- Make calls idempotent where you can, so a retry or replay doesn't repeat a side effect.
- Write specific tool names and descriptions. They become function names and doc comments in the generated API.
- Check permissions on every call in your own code, whoever the caller is.
The Model Context Protocol itself doesn't need to change for any of this. These are design choices on the server side.
The security trade-off
Code Mode means running code the model wrote, and a prompt injection can steer that code. Anthropic's MCP post says plainly that this needs a sandbox, resource limits and monitoring, which direct tool calls don't. The hosted versions from Anthropic and OpenAI handle the sandbox for you. If you host it yourself, you need real isolation. Our E2B vs Daytona vs Modal comparison covers sandbox options like E2B and Daytona.
Isolates aren't magic. In August 2026, Check Point Research disclosed five vulnerabilities in workerd, the runtime under Cloudflare Workers and Code Mode, and Cloudflare rated two of them Critical. One exploit chain started from a prompt injection and escaped the Code Mode sandbox on a self-hosted build. Cloudflare fixed its managed platform, and self-hosted workerd needs v1.20260619.1 or later. The lesson is to block network access by default, keep secrets on the host side (Cloudflare's bindings and OpenAPI wrapper do this), keep runtimes patched, and require approval for writes.
Who should use Code Mode
- Agent builders on the Claude or OpenAI APIs with many tools or large tool outputs: try programmatic tool calling on one bounded stage and measure it.
- Teams on Cloudflare Workers:
@cloudflare/codemodegives you the sandbox, MCP wrapping and approval flows in one package. - Coding-agent users: Pi uses it for MCP today. Codex's version is still experimental. For a wider view of the harnesses, see Claude Code vs Codex vs Gemini CLI.
- Skip it if your agent makes a few small calls per task. Plain tool calling is simpler and may be cheaper.
Code Mode also fits well with agent skills. Anthropic suggests saving a working script with a SKILL.md file so the agent can reuse it later.
FAQ
What is Code Mode in MCP?
It's a way of using MCP servers where the agent sees their tools as a typed code API and writes a script that calls them, instead of calling each tool directly. The script runs in a sandbox, and only its output returns to the model.
Is Code Mode the same as programmatic tool calling?
They're the same idea under different names. Cloudflare calls it Code Mode. Anthropic and OpenAI call their hosted API features programmatic tool calling. Claude writes Python, while OpenAI's runtime and Cloudflare's package use JavaScript.
Does Code Mode replace MCP?
No. MCP still handles connecting to servers, listing their tools and authorization. Code Mode only changes how the agent calls those tools.
How do I turn on code mode in Codex?
Set enabled = true under [features.code_mode] in your Codex config file. OpenAI's configuration reference says the feature is under development and off by default, so expect changes.
Does Code Mode always save tokens?
No. Anthropic's tests showed large savings with many tools and big results, but about 8% higher cost on a benchmark where each turn made only one or two sequential calls. Measure on your own workload.
Is Code Mode safe?
It's as safe as the sandbox it runs in. Treat model-written code as untrusted: block network access, keep credentials outside the sandbox, patch the runtime, and require approval for actions that change data.



