Docker Agent is Docker's open-source runtime for building and running AI agents from a config file instead of code. You describe each agent in YAML (or HCL): its model, its instructions, its tools and which other agents it can delegate to. Then docker agent run agent.yaml starts it in a terminal UI. The same project was called cagent in earlier Docker Desktop versions, and it's Apache-2.0 licensed on GitHub.

This guide explains what Docker Agent actually does, how to install it and write a first config, how it handles tools, MCP and multi-agent teams, the safety controls you should know about, and how it compares with code-first frameworks. Facts are checked against Docker's documentation and the v1.149.0 release (October 7, 2026).

What is Docker Agent?

Docker's official Docker Agent page describes it as a framework for building teams of specialized agents that collaborate, run from your terminal with any LLM provider. In practice it is three things in one binary:

  • A runtime. It runs the agent loop for you: send the prompt to the model, execute the tool calls, feed results back, repeat.
  • A declarative config format. Agents, models, tools and delegation rules live in a versionable YAML or HCL file. There is no Python or TypeScript glue code to write.
  • A distribution format. Agent configs can be pushed to and pulled from any OCI registry, such as Docker Hub or GitHub Container Registry, the same way container images are.

It ships as a docker CLI plugin, so you run it as docker agent, but it also works as a standalone docker-agent binary. It isn't Gordon, Docker's built-in assistant (docker ai); Docker's docs draw that line explicitly.

Is cagent the same as Docker Agent?

Yes. Docker's docs say the feature was called cagent in Docker Desktop 4.49 through 4.62, and it ships as Docker Agent from Docker Desktop 4.63 onward. The old github.com/docker/cagent URL now redirects to docker/docker-agent. Older tutorials that mention cagent describe the same tool under its former name.

How a Docker Agent config works

Every config has a root agent, the one that receives your messages. In a single-agent setup it's the only one. In a team, it acts as the coordinator. A small two-agent team looks like this:

agents:
  root:
    model: anthropic/claude-sonnet-4-5
    description: Bug investigator
    instruction: |
      Find the root cause of the error the user pastes.
      Hand the fix to the fixer agent.
    sub_agents: [fixer]
    toolsets:
      - type: filesystem
      - type: think

  fixer:
    model: openai/gpt-5-mini
    description: Writes minimal, tested fixes
    instruction: Make the smallest change that fixes the bug and add a test.
    toolsets:
      - type: filesystem
      - type: shell

The key fields, per the agents reference:

  • model uses a provider/model reference. Built-in providers include OpenAI, Anthropic, Google Gemini, AWS Bedrock, Mistral, xAI, OpenRouter, Docker Model Runner (dmr) and local OpenAI-compatible servers such as Ollama.
  • description is what other agents read when deciding whom to delegate to, so write it for them.
  • instruction is the system prompt.
  • toolsets lists built-in tools and MCP servers.
  • sub_agents names the agents this one can hand work to.
  • Optional extras include fallback models (with retries and a cooldown after rate limits), max_iterations to cap tool loops, named commands you can call as /name, and skills to load SKILL.md skills (see our agent skills guide).

Each agent has its own model and its own context. Agents don't share knowledge automatically; the parent passes a task description and gets a result back.

How to install and run Docker Agent

  1. Install it. Docker Desktop 4.63 and later include it. Run docker agent version to check. Without Desktop, use brew install docker-agent, winget install Docker.Agent, or a binary from the releases page. Copy the binary into ~/.docker/cli-plugins if you want the docker agent form.
  2. Give it a model. Export a provider key such as ANTHROPIC_API_KEY or OPENAI_API_KEY. To avoid API costs, point the config at a local model through Docker Model Runner or Ollama instead.
  3. Run something. docker agent run with no file starts a built-in general assistant (or a docker-agent.yaml in the current folder, if there is one). docker agent new builds a config through an interactive wizard. docker agent run agent.yaml runs your own.
  4. Go headless when you need to. docker agent run --exec agent.yaml "Create a Dockerfile for a Node.js app" runs one task and exits, and you can pipe a log file into it.

The quick start also has a scripted tour you can start with docker agent getting-started.

Tools and MCP in Docker Agent

Built-in toolsets cover the usual building blocks: filesystem, shell, think, todo, memory (backed by a local database file), fetch, RAG with BM25, embedding or hybrid search, and task delegation. Side-effecting tools such as shell commands and file writes ask for confirmation by default.

For everything else there's the Model Context Protocol. Docker Agent connects to MCP servers in three ways:

  • Docker MCP, the recommended path: ref: docker:duckduckgo runs a server from Docker's MCP catalog in a container through the MCP Gateway. This needs Docker Desktop.
  • Local stdio servers started as subprocesses, such as one you wrote yourself with our Python MCP server tutorial.
  • Remote servers over Streamable HTTP or SSE.

A newer mcp_catalog toolset lets the agent search and switch on servers from a curated catalog subset during a run, with allow and block lists, so the prompt isn't flooded with tool definitions it never uses. Docker Agent also has a Code Mode feature; our Code Mode explainer covers why that pattern saves tokens.

Multi-agent teams, handoffs and coding harnesses

Docker Agent supports two team patterns, described on its multi-agent page:

  • Delegation (sub_agents). The parent calls a built-in transfer_task tool with a task and the expected output. The child works in its own sub-session and returns a result, and the parent continues. Delegation is auto-approved.
  • Handoffs (handoffs). The whole conversation moves to another agent, which sees the full history. Agents can form loops, which suits pipelines and routing.

There are also background agents for independent tasks that can run at the same time.

The most interesting addition for coding work is harnesses. An agent with a harness: block hands its work to an external coding CLI instead of calling a model API: Claude Code, Codex, opencode or pi. The harness docs frame it as letting the CLI drive the coding loop while Docker Agent adds orchestration, hooks, permissions and distribution. So a planner agent could split a job and pass the coding pieces to Claude Code or Codex. The CLI must be installed and on your PATH, and Claude Code must be logged in. If you're choosing between those CLIs, see our Claude Code vs Codex vs Gemini CLI comparison.

Safety: permissions, sandbox mode and telemetry

An agent with a shell tool can run commands on your machine, so these controls matter.

Permissions and safety modes. Permission rules mark tools as allowed, ask-first or denied, evaluated deny first, then allow, then ask. Unmatched calls fall back to a safety mode: strict asks about everything, balanced runs safe calls (such as ls or git status) silently and asks about the rest, restricted silently denies anything not known to be safe, and autonomous (the old --yolo flag) runs everything. Docker says plainly that restricted is defense in depth, not a security boundary.

Sandbox mode. For real isolation, docker agent run --sandbox agent.yaml runs the agent inside a Docker Sandboxes microVM. Your working directory is mounted read-write and Docker Agent's own config folder read-only. The agent can still change files in your project, but it can only reach what is mounted, not the rest of your machine. A --cloud option runs on Docker-managed compute instead. It doesn't upload your local folder, config or API keys, so it needs a built-in or published agent and secrets set up on the cloud side. Docker says the sbx CLI and local sandboxes are free, including for commercial work, while cloud compute is pay-as-you-go. For other hosted sandboxes, see E2B vs Daytona vs Modal.

Telemetry. Usage telemetry is on by default. Docker's telemetry page says command events include positional arguments, which for run and exec can contain your prompt and file paths. Set TELEMETRY_ENABLED=false before running anything sensitive.

Sharing, serving and testing agents

  • Share. docker agent share push ./agent.yaml docker.io/you/my-agent:latest publishes a config, and anyone can run it with docker agent run you/my-agent:latest. A --key option signs the artifact so pullers can verify it wasn't altered.
  • Serve over MCP. docker agent serve mcp ./agent.yaml exposes your agent as an MCP tool for Claude Desktop, Claude Code, Cursor and other clients, over stdio or streaming HTTP.
  • Serve over A2A or HTTP. docker agent serve a2a speaks Google's Agent2Agent protocol (Docker calls that support early), and docker agent serve api exposes a REST API with streaming. Our MCP vs A2A guide explains the difference.
  • Evaluate. docker agent eval replays recorded sessions in clean containers and scores tool calls and answers.

Docker Agent vs code-first agent frameworks

ItemDocker AgentLangGraphCrewAI
How you define agentsYAML or HCL configPython or JS code (graphs)Python code (crews, flows)
RuntimeSingle Go binary or Docker CLI pluginLibrary in your appLibrary in your app
SharingOCI registry push and pullYour own packagingYour own packaging
LicenseApache-2.0MITMIT

The trade-off is control. LangGraph and CrewAI let you write arbitrary logic between steps and embed agents inside an application. Docker Agent gives you a working agent, tools and team setup with no code, and features like sandboxing and registry sharing built in, but complex custom behavior has to fit its config schema, hooks or Go SDK. Our LangGraph vs CrewAI vs OpenAI Agents SDK comparison covers the code-first side.

Who Docker Agent is for

Choose it if you already use Docker, want a terminal agent you can configure without writing code, want to mix models from several providers in one team, or need to share a standard agent across a team through a registry.

Look elsewhere if you're embedding agents in a product and need fine control over state and branching, or you want a polished IDE experience. A dedicated coding agent will feel smoother for day-to-day editing.

Pros: open source, provider-agnostic, containerized MCP tools, real sandbox option, frequent releases.
Cons: the docs on GitHub Pages track the main branch and can describe unreleased features, telemetry is opt-out, and at least one commenter in the October 2026 Hacker News discussion called it more brittle than other harnesses they had tried.

FAQ

What is Docker Agent?

Docker Agent is an open-source tool from Docker that runs AI agents and agent teams defined in a YAML or HCL file. It handles the model calls, tools, MCP servers and delegation, and it ships as the docker agent CLI plugin.

Is cagent the same as Docker Agent?

Yes. It was called cagent in Docker Desktop 4.49 to 4.62 and was renamed Docker Agent from 4.63. The GitHub repository moved from docker/cagent to docker/docker-agent.

Is Docker Agent free?

The tool is free and open source under Apache-2.0. You pay your model provider for API usage unless you run local models. Docker Sandboxes are free locally, and cloud sandboxes are billed pay-as-you-go.

Do I need Docker Desktop to use Docker Agent?

No. You can install the binary with Homebrew, winget or a GitHub release. Docker Desktop is needed for containerized docker: MCP tools and Docker Model Runner.

Can Docker Agent use local models?

Yes. Use Docker Model Runner models with the dmr/ prefix, or point it at Ollama, vLLM or another OpenAI-compatible server. No API key is needed for those.

Is it safe to give Docker Agent shell access?

It asks before side-effecting tool calls by default, and you can tighten that with permission rules and safety modes. For real isolation, run it with --sandbox so it works inside a Docker Sandboxes microVM.