Back to Browse

Codex Subagent MCP Server

Developer ToolsLow Risk10.0MCP RegistryLocal
Free

Server data from the Official MCP Registry

Delegate coding tasks to the OpenAI Codex CLI, installed separately, with explicit model and effort.

About

Delegate coding tasks to the OpenAI Codex CLI, installed separately, with explicit model and effort.

Security Report

10.0
Low Risk10.0Low Risk

Valid MCP server (1 strong, 1 medium validity signals). No known CVEs in dependencies. Package registry verified. Imported from the Official MCP Registry.

6 files analyzed · 1 issue found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

env_vars

Check that this permission is expected for this type of plugin.

Shell Command Execution

Runs commands on your machine. Be cautious — only use if you trust this plugin.

file_system

Check that this permission is expected for this type of plugin.

What You'll Need

Set these up before or after installing:

Model used when a call names none; without it the server asks which model to use.Optional

Environment variable: CODEX_SUBAGENT_DEFAULT_MODEL

Reasoning effort used when a call specifies none.Optional

Environment variable: CODEX_SUBAGENT_DEFAULT_EFFORT

Comma-separated allow-list of model slugs; anything else is refused.Optional

Environment variable: CODEX_SUBAGENT_ALLOWED_MODELS

Sandbox used when a call specifies none. Defaults to read-only and cannot exceed the ceiling.Optional

Environment variable: CODEX_SUBAGENT_DEFAULT_SANDBOX

Ceiling on what a delegation may do. Defaults to workspace-write; danger-full-access must be set explicitly.Optional

Environment variable: CODEX_SUBAGENT_MAX_SANDBOX

Ceiling on reasoning effort.Optional

Environment variable: CODEX_SUBAGENT_MAX_EFFORT

Path to the Codex executable, if it is not codex on PATH. On Windows it must be codex.exe, not a .cmd shim.Optional

Environment variable: CODEX_BIN

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-parisbs-codex-subagent-mcp": {
      "env": {
        "CODEX_BIN": "your-codex-bin-here",
        "CODEX_SUBAGENT_MAX_EFFORT": "your-codex-subagent-max-effort-here",
        "CODEX_SUBAGENT_MAX_SANDBOX": "your-codex-subagent-max-sandbox-here",
        "CODEX_SUBAGENT_DEFAULT_MODEL": "your-codex-subagent-default-model-here",
        "CODEX_SUBAGENT_ALLOWED_MODELS": "your-codex-subagent-allowed-models-here",
        "CODEX_SUBAGENT_DEFAULT_EFFORT": "your-codex-subagent-default-effort-here",
        "CODEX_SUBAGENT_DEFAULT_SANDBOX": "your-codex-subagent-default-sandbox-here"
      },
      "args": [
        "-y",
        "codex-subagent-mcp"
      ],
      "command": "npx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

codex-subagent-mcp

CI License: MIT

An MCP server that lets Claude Code delegate coding tasks to OpenAI's Codex CLI running on the same machine — multi-model orchestration, locally, with the model and reasoning depth chosen per task.

Claude stays the orchestrator. Codex becomes a subagent it can call.

This runs another agent on your machine. Codex can read local files and run commands with the permissions you grant it. Read the security model before enabling writes or unsandboxed runs.

An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic.

Why this exists

A single model doing everything has three recurring problems, and delegation solves each one:

Your context window is finite. Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files.

One model has one set of blind spots. A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice.

Not every task deserves the same reasoning budget. Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the model sets raw capability, the reasoning effort sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs.

Everything stays on your machine. The server drives the Codex CLI you already have installed and holds no credentials of its own.

See How it compares for a versioned comparison with other Codex MCP servers.

Requirements

  • Node.js 22 or newer.
  • The Codex CLI, installed, on PATH, and signed in.

You do not have to check this by hand. Run the codex_doctor tool — or just ask Claude to — and it reports what is missing and the exact commands for your platform. Every tool that reaches the CLI runs the same check first, so you never get a bare spawn ENOENT. Nothing is ever installed on your behalf.

If you do not have the Codex CLI yet, install it without npm:

# macOS — recommended
brew install --cask codex
# macOS / Linux — standalone installer
curl -fsSL https://chatgpt.com/codex/install.sh | sh
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"

Then run codex once to sign in, and confirm with codex login status. On Windows, open a new terminal first so the updated PATH is picked up.

Codex's sandbox depends on the platform, so two notes from OpenAI's documentation:

  • Linux and WSL2 — Codex sandboxes commands with bubblewrap. Install it with your package manager before the first delegation. Without it Codex falls back to a bundled helper that needs unprivileged user namespaces, which some distributions restrict. See sandboxing.
  • Windows — Codex runs natively, without WSL, and uses its own Windows sandbox. Windows 11 is recommended; Windows 10 version 1809 or newer is the practical minimum. See the Windows sandbox documentation.

The npm package is not the program: bin/codex.js is a Node wrapper that spawns the real Rust binary. That has two consequences.

Every invocation pays a Node startup. Measured on macOS: about 80 ms through the wrapper against about 20 ms calling the binary directly. This server spawns the CLI once per tool call, so the cost recurs — though it is still noise next to a delegation that runs for seconds.

A global npm install lives inside the active Node version. Under a version manager such as nvm it lands in ~/.nvm/versions/node/<version>/lib/node_modules, so switching Node versions takes codex off PATH until you reinstall it. This is the bigger problem in practice.

On Windows there is a third, harder consequence: a global npm install produces a codex.cmd batch shim, which cannot be launched without a command shell — and this server never uses one. It detects that case and says so, but the installer avoids it entirely. See ADR 11.

Switching is two commands, and your sign-in survives because credentials live in Codex's home directory — ~/.codex, or %USERPROFILE%\.codex on native Windows — not in the npm package:

# macOS
npm uninstall -g @openai/codex && brew install --cask codex
# Linux
npm uninstall -g @openai/codex && curl -fsSL https://chatgpt.com/codex/install.sh | sh
# Windows (PowerShell) — two lines, because Windows PowerShell rejects `&&`
npm uninstall -g @openai/codex
powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"

Install

claude mcp add codex-subagent -- npx -y codex-subagent-mcp

That works in both the Claude Code CLI and the desktop app; they share the same configuration.

Clients that browse the official MCP Registry list version 0.4.0 under io.github.parisbs/codex-subagent-mcp (verified 2026-09-30). That entry installs the same codex-subagent-mcp npm package.

Global install, if you prefer not to go through npx:

npm install -g codex-subagent-mcp

Then point Claude Code at the codex-subagent binary.

Claude Desktop has no equivalent command. Add an entry to mcpServers in claude_desktop_config.json and restart the app:

{
  "mcpServers": {
    "codex-subagent": {
      "command": "npx",
      "args": ["-y", "codex-subagent-mcp"]
    }
  }
}

Settings → Developer → Edit Config opens the file and creates it if it does not exist. Its documented locations are ~/Library/Application Support/Claude/claude_desktop_config.json on macOS and %APPDATA%\Claude\claude_desktop_config.json on Windows. The Linux desktop app is in beta, and its documentation does not say where the file lives. If the server does not appear after a restart, the MCP logs are in ~/Library/Logs/Claude on macOS and %APPDATA%\Claude\logs on Windows. See Connect to local MCP servers.

From a clone, for development:

git clone https://github.com/parisbs/codex-subagent-mcp.git
cd codex-subagent-mcp && npm ci && npm run build

The repository ships a .mcp.json, so running Claude Code from the project root picks the server up.

Set CODEX_BIN if your Codex executable is not called codex or is not on PATH. On Windows, point it at the real codex.exe: a .cmd or .bat shim is refused rather than run through a shell.

First steps

Ask Claude to check the installation:

Check that the Codex subagent is set up correctly.

You should see status: ok, a version, and signed in: yes. Then see what you can delegate to:

What Codex models are available, and what are they each good for?

Then try a real one. With the built-in default this is read-only, so Codex investigates and reports without touching anything:

Have Codex look at this repository and explain how the build is wired together.

If you have not set a default model, the tool descriptions tell Claude to call codex_recommend first, to give you the suggested model and effort in the same message in which it says it is going to delegate, and then to call codex_delegate with both values explicit. Those descriptions are guidance to a model rather than a rule it cannot break, and a delegation that reaches this server with no model at all is refused rather than guessed at. The recommendation is advice, not a decision on your behalf. To skip that step on later delegations, set CODEX_SUBAGENT_DEFAULT_MODEL once as shown in Choosing a model.

Using it

Delegations run read-only by default: Codex investigates and reports, but cannot modify files. Letting it write is a deliberate call argument or a default you set in the server environment.

Write a bounded delegation

A delegation gets expensive when repeated commands keep adding output to the context carried into later requests. Name the exact question, likely files, stopping condition and evidence the answer must contain; choose higher effort for ambiguity rather than by habit. Writing a delegation gives the measured cost model, ranked rules and weak-versus-strong examples using the real tool parameters. Background goes in context and a persona or extra rules in system_instructions, both layered on the built-in quality contract; see the tool reference.

The examples below are the four situations where delegating beats doing it in the main conversation. Each one has been run against the real Codex CLI — writing them is how two defects in this server were found and fixed.

Get a second opinion from a different model family

The value here is not a second run — it is a different set of blind spots.

Ask Codex to review src/server.ts for correctness problems, focusing on error paths. Use a high reasoning effort and tell it to report each finding with the line and why it matters.

Claude picks the model, passes your file as the focus, and returns the findings. This is how the terminate() defect in this repository's own runner was found: a delegated review spotted that two code paths could each arm a timer while only one was ever cleared.

Investigate without spending your context

Forty files go into the delegation; one answer comes back.

Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments handed to the Codex CLI, and report just the call chain.

Codex runs its own searches and reads whatever it needs. Your conversation receives the conclusion, not the search results.

Run long work in the background while you keep going

Kick off a Codex run in the background that writes unit tests for src/jobs.ts, then keep helping me with the API layer.

You get a job_id immediately. Ask for the status whenever you want, and read the result when it is done. Up to eight can run at once.

Buy deep reasoning for one hard problem

Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not.

This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at maximum reasoning effort, give it test/runner.test.ts and the CI log, and tell it not to change anything — I want the diagnosis first.

Keep the thread going

Follow-ups reuse the context Codex already has, so they cost a fraction of the original:

Ask Codex to expand on its second finding.

The follow-up runs on the same model, effort and directory as the original. Codex itself does not keep those when a session resumes, so the server restates them for every thread it started.

Let it write, when you mean it

Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree so my working tree stays clean.

That last clause matters: use_worktree sends the run's own edits to a managed git worktree under ~/.codex/worktrees/ instead of your checkout, and the result lists the files it touched with the path where each landed — up to a thousand distinct files, after which it says how many it left out. It is not a second sandbox: what a run may write outside that worktree is still decided by the sandbox and by add_dirs. Worktrees rely on an experimental Codex feature, which the server turns on for that invocation only — it never changes your Codex configuration.

Whatever the sandbox, a delegation that writes reports what it wrote:

Files changed (2):
- [edit] src/codex/runner.ts
- [add] test/runner.test.ts

Before you start enabling writes as a habit, read the next section. It is short.

Safety

This server runs another program on your machine, so it is worth two minutes before you enable writes.

What protects you

Delegations are read-only by default. Writing requires an explicit sandbox: "workspace-write" or a user-set default, and use_worktree sends the run's edits to a managed git worktree instead of your checkout — while the sandbox and add_dirs, not the worktree, are what bound where it can write at all. Unsandboxed runs are unavailable unless you explicitly opt into that ceiling.

The confinement is not a promise from the model — it is the operating system's own sandbox: Seatbelt on macOS, bubblewrap on Linux and WSL2, and a native sandbox on Windows. The table was measured on macOS against Codex CLI 0.154.0. Linux and Windows have not been measured here, and on Windows OpenAI's documentation notes that sandboxed commands can fail to read some directories, so reads may be stricter there:

read-onlyworkspace-writedanger-full-access
Write inside the working directorynoyesyes
Write outside it (your home)nonoyes
Network accessnonoyes
Read outside the working directoryyesyesyes

There is also no shell anywhere in the path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert.

What does not protect you

Reads are not confined. That last row is not a typo. Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel.

A prompt is untrusted input, and Codex acts on it. This is prompt injection, and it is the risk that matters here. If you build a delegation from content you did not write — an issue body, a web page, a log, a file from someone else's repository — that content can carry instructions. With workspace-write it can direct Codex to modify your repository; even read-only it can direct Codex to read something sensitive and put it in the answer. The sandbox bounds where Codex can write. It does not judge what it should write, or why it was asked.

The result is not sanitised. What comes back is text from a model that just read your files. Treat it as data, not as instructions.

Reducing the risk

  • Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation.
  • If you never want writes from this server, cap it: CODEX_SUBAGENT_MAX_SANDBOX=read-only. A ceiling cannot be argued past by anything in the conversation, which is what makes it different from a default. Register it outside the repository (Claude Code's default local scope, --scope user, or Claude Desktop's config), not in a project .mcp.json that a write-enabled delegation could edit. See Configuration.
  • When you do enable writes, add use_worktree so changes land somewhere you can inspect before they touch your branch.
  • Do not assemble delegation prompts from untrusted content when you intend to act on the answer.
  • If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it.

SECURITY.md has the full threat model, what a deny_read policy could add, and how to report a vulnerability.

Staying in control

Claude decides when to delegate, and every delegation sends its prompt to OpenAI and spends your Codex usage — in any conversation where the server is available, not only programming ones. The defaults are safe, and you can tighten them in layers:

  • Your client's permission prompt. Let the inspection tools run freely, and keep confirming codex_delegate and codex_follow_up, the two that spend usage.
  • Ceilings on the server, such as CODEX_SUBAGENT_MAX_SANDBOX and CODEX_SUBAGENT_MAX_EFFORT, which no argument can get past.
  • A version range such as codex-subagent-mcp@^0.4.0, so new behaviour arrives when you choose.
  • Your own rules in CLAUDE.md, for when Claude should delegate at all.

docs/CONTROL.md shows how to set each one, what the server already does on its own, and what no setting can guarantee.

Choosing a model

Read live from your installed CLI, so this list tracks whatever you have. As of Codex CLI 0.159.2:

SlugPositioningReasoning effortsDefault
gpt-6.1-solLatest workhorse for coding and everyday worklow … ultralow
gpt-6-astraFrontier intelligence for the most demanding worklow … ultralow
gpt-6-solPrevious generation workhorselow … ultramedium
gpt-6-lunaFast and affordable, for easier taskslow … maxmedium
gpt-5.6-solOlder generation workhorselow … ultralow
gpt-5.6-terraOlder balanced model for straightforward worklow … ultramedium
gpt-5.6-lunaOlder fast and efficient modellow … maxmedium
gpt-5.5Legacy coding modellow … xhighmedium

Model and reasoning effort are independent. The model sets raw capability; the effort — low, medium, high, xhigh, max, ultra — sets how long it deliberates before acting. ultra additionally delegates subtasks automatically.

The server does not choose for you. Which model a task deserves depends on your budget and on how costly a wrong answer is, and a regular expression over a prompt cannot know either. Ask for a delegation without naming a model and it refuses — but the refusal carries the recommendation it would have made, so you decide in one more exchange instead of paying for a guess.

If you would rather not be asked, set a default once and it stops asking:

claude mcp add codex-subagent -e CODEX_SUBAGENT_DEFAULT_MODEL=gpt-5.6-terra -- npx -y codex-subagent-mcp

For advice rather than a decision, ask:

Which Codex model should handle migrating this repo's tests to vitest?

That routes mechanical edits to the fast model at low, everyday work to the balanced one at medium, multi-file migrations to the agentic workhorse at high, and hard reasoning problems to the most capable model at xhigh or ultra. It is a suggestion you can ignore, and it stays within the model allow-list and effort ceiling you configure. An effort the chosen model does not support is adjusted to the closest level it does, with a note saying so.

Configuration

Everything is optional, and set through environment variables on the MCP server:

VariableEffect
CODEX_SUBAGENT_DEFAULT_MODELStops the server asking which model to use.
CODEX_SUBAGENT_DEFAULT_EFFORTReasoning effort when a call specifies none.
CODEX_SUBAGENT_ALLOWED_MODELSComma-separated allow-list. Anything else is refused.
CODEX_SUBAGENT_DEFAULT_SANDBOXSandbox when a call specifies none. Defaults to read-only and cannot exceed the ceiling.
CODEX_SUBAGENT_MAX_SANDBOXCeiling on what a delegation may do. Defaults to workspace-write; it must be set to danger-full-access explicitly before unsandboxed calls are allowed.
CODEX_SUBAGENT_MAX_EFFORTCeiling on reasoning effort. Useful for keeping ultra off the table. A call above it is lowered to a level the model supports, or refused if the model has none that low.
CODEX_BINPath to the Codex executable, if it is not codex on PATH. On Windows it must be codex.exe, not a .cmd shim.

The sandbox settings express policy you choose outside the repository: a default saves repeated arguments, while the ceiling is the boundary no call can cross. See Safety for why the built-in ceiling stops at workspace-write.

Your own escalation rules belong in your CLAUDE.md, in plain language, where Claude applies them with actual understanding and they stay yours. See ADR 12 for why they are not built into this server, and ADR 14 for the sandbox policy split.

Tools

ToolWhat it does
codex_doctorCheck the Codex CLI installation and report how to fix it.
list_codex_modelsList available models and their reasoning-effort levels.
codex_recommendSuggest a model and effort for a described task.
codex_delegateRun a task, blocking or in the background, optionally returning JSON that matches a schema.
codex_follow_upContinue a previous delegation using its thread_id, with or without a schema.
codex_job_statusCheck a background delegation.
codex_job_resultRead a finished background delegation's output.
codex_job_cancelStop a running background delegation.

Full parameter reference: docs/TOOLS.md.

FAQ

Does this cost money? It uses your existing Codex quota, the same as running codex yourself. This server adds nothing. Higher reasoning efforts consume more; codex_recommend exists partly so you do not spend ultra on work that low would have handled.

Can it modify my files? Not by default. Delegations run read-only unless you explicitly ask for write access, and use_worktree keeps even those changes out of your working tree.

Why drive the CLI instead of calling the OpenAI API? Delegated coding is not a single completion — it is an agentic loop with a sandbox, an approval model, session persistence and project instruction files. All of that lives in the Codex client, not in the model endpoint. See ADR 1.

Do I need Claude Code, or does Claude Desktop work? Either. Claude Code gets a one-line install; Claude Desktop needs a manual config entry.

It says Codex is not installed, but codex works in my terminal. Most likely Windows with a global npm install, which produces a codex.cmd batch shim that cannot be launched without a command shell. codex_doctor reports this as unsupported-shim and offers two fixes. On macOS and Linux, check whether a Node version manager moved codex off PATH.

Does it work on Windows and Linux? CI builds, tests and starts the server on Windows, macOS and Linux on every change, and checks that the Codex CLI is resolved correctly on each. A real delegation has only been verified on macOS — the CI runners have no Codex installation or credentials. Reports from Windows and Linux are welcome. On Windows, install the Codex CLI with the PowerShell installer rather than npm; on Linux, install bubblewrap for Codex's sandbox. See Requirements.

Can Codex read files outside the directory I point it at? Yes, in every sandbox mode — the sandbox restricts writes and network access, not reads. See Safety for what that means in practice and what to do about it.

Where do worktree changes end up? Under ~/.codex/worktrees/, and the delegation result gives you the full path of each file it touched, up to a thousand distinct files. The server does not clean those worktrees up: they may hold work you have not applied yet.

Documentation

  • docs/DELEGATING.md — how to scope a delegation, with measured costs and worked prompts.
  • docs/TOOLS.md — every tool and parameter.
  • docs/COMPARISON.md — versioned comparisons with other Codex MCP servers.
  • docs/adr/ — why the design is what it is, decision by decision.
  • docs/ROADMAP.md — what is planned, and what is deliberately out of scope.
  • Issues — what is actually open right now.
  • docs/VERSIONING.md — what counts as a breaking change here.
  • CHANGELOG.md — what changed in each release, and the Codex CLI version it was verified against.
  • CONTRIBUTING.md — setup, and the rules that are not negotiable.

Disclaimer

Not an official product. This is an independent, community project. It is not affiliated with, endorsed by, sponsored by or supported by OpenAI or Anthropic. "Codex", "ChatGPT" and "OpenAI" are trademarks of OpenAI; "Claude" and "Claude Code" are trademarks of Anthropic. They are used here only to describe what this software interoperates with, which is nominative use — no claim is made to any of them. Neither company is responsible for this software, and problems with it should be reported here rather than to them.

No warranty. The software is provided "as is", without warranty of any kind, as stated in LICENSE. You use it at your own risk.

It runs an agent on your machine. This server spawns the Codex CLI as a child process. Depending on the sandbox you allow, that process can read your files, run shell commands and modify your working tree. Read Safety before enabling writes, and review what a delegation did rather than assuming it did what you asked.

It spends your quota. Delegations consume your own OpenAI Codex usage, at whatever rate your account is billed. Higher reasoning efforts consume more, and ultra delegates subtasks of its own. This project has no visibility into that cost and does not cap it beyond the limits you configure yourself.

License

MIT. See LICENSE.

Reviews

No reviews yet

Be the first to review this server!