Skip to main content
Glama

Codex Octopus

One brain, many arms.

An MCP server that wraps the OpenAI Codex SDK, letting you run multiple specialized Codex agents — each with its own model, sandbox, effort, and personality — from any MCP client.

Why

Codex is powerful. But one instance does everything the same way. Sometimes you want a strict code reviewer in read-only sandbox. A test writer with workspace-write access. A cheap quick helper on minimal effort. A deep thinker on xhigh.

Codex Octopus lets you spin up as many of these as you need. Same binary, different configurations. Each one shows up as a separate tool in your MCP client.

Related MCP server: Codex MCP Server

Prerequisites

  • Node.js >= 18

  • Codex CLI — the Codex SDK spawns the Codex CLI under the hood, so you need it installed (@openai/codex)

  • OpenAI API key (CODEX_API_KEY env var) or inherited from parent process

Install

npm install codex-octopus

Or use npx directly in your .mcp.json (see Quick Start below).

Quick Start

Add to your .mcp.json:

{
  "mcpServers": {
    "codex": {
      "command": "npx",
      "args": ["codex-octopus@latest"],
      "env": {
        "CODEX_SANDBOX_MODE": "workspace-write",
        "CODEX_APPROVAL_POLICY": "never"
      }
    }
  }
}

This gives you two tools: codex and codex_reply. That's it — you have Codex as a tool.

Multiple Agents

The real power is running several instances with different configurations:

{
  "mcpServers": {
    "code-reviewer": {
      "command": "npx",
      "args": ["codex-octopus@latest"],
      "env": {
        "CODEX_TOOL_NAME": "code_reviewer",
        "CODEX_SERVER_NAME": "code-reviewer",
        "CODEX_DESCRIPTION": "Strict code reviewer. Read-only sandbox.",
        "CODEX_MODEL": "o3",
        "CODEX_SANDBOX_MODE": "read-only",
        "CODEX_APPEND_INSTRUCTIONS": "You are a strict code reviewer. Report real bugs, not style preferences.",
        "CODEX_EFFORT": "high"
      }
    },
    "test-writer": {
      "command": "npx",
      "args": ["codex-octopus@latest"],
      "env": {
        "CODEX_TOOL_NAME": "test_writer",
        "CODEX_SERVER_NAME": "test-writer",
        "CODEX_DESCRIPTION": "Writes thorough tests with edge case coverage.",
        "CODEX_MODEL": "gpt-5-codex",
        "CODEX_SANDBOX_MODE": "workspace-write",
        "CODEX_APPEND_INSTRUCTIONS": "Write tests first. Cover edge cases. TDD."
      }
    },
    "quick-qa": {
      "command": "npx",
      "args": ["codex-octopus@latest"],
      "env": {
        "CODEX_TOOL_NAME": "quick_qa",
        "CODEX_SERVER_NAME": "quick-qa",
        "CODEX_DESCRIPTION": "Fast answers to quick coding questions.",
        "CODEX_EFFORT": "minimal"
      }
    }
  }
}

Your MCP client now sees three distinct tools — code_reviewer, test_writer, quick_qa — each purpose-built.

Agent Factory

Don't want to write configs by hand? Add a factory instance:

{
  "mcpServers": {
    "agent-factory": {
      "command": "npx",
      "args": ["codex-octopus@latest"],
      "env": {
        "CODEX_FACTORY_ONLY": "true",
        "CODEX_SERVER_NAME": "agent-factory"
      }
    }
  }
}

This exposes a single create_codex_mcp tool — an interactive wizard. Tell it what you want ("a strict code reviewer with read-only sandbox") and it generates the .mcp.json entry for you.

Tools

Each non-factory instance exposes:

Tool

Purpose

<name>

Send a task to the agent, get a response + thread_id

<name>_reply

Continue a previous conversation by thread_id

Per-invocation parameters (override server defaults):

Parameter

Description

prompt

The task or question (required)

cwd

Working directory override

model

Model override

additionalDirs

Extra directories the agent can access

effort

Reasoning effort (minimal to xhigh)

sandboxMode

Sandbox override (can only tighten, never loosen)

approvalPolicy

Approval override (can only tighten, never loosen)

networkAccess

Enable network access from sandbox

webSearchMode

Web search: disabled, cached, live

instructions

Additional instructions (prepended to prompt)

Configuration

All configuration is via environment variables in .mcp.json. Every env var is optional.

Identity

Env Var

Description

Default

CODEX_TOOL_NAME

Tool name prefix (<name> and <name>_reply)

codex

CODEX_DESCRIPTION

Tool description shown to the host AI

generic

CODEX_SERVER_NAME

MCP server name in protocol handshake

codex-octopus

CODEX_FACTORY_ONLY

Only expose the factory wizard tool

false

Agent

Env Var

Description

Default

CODEX_MODEL

Model (gpt-5-codex, o3, codex-1, etc.)

SDK default

CODEX_CWD

Working directory

process.cwd()

CODEX_SANDBOX_MODE

read-only, workspace-write, danger-full-access

read-only

CODEX_APPROVAL_POLICY

never, on-failure, on-request, untrusted

on-failure

CODEX_EFFORT

minimal, low, medium, high, xhigh

SDK default

CODEX_ADDITIONAL_DIRS

Extra directories (comma-separated)

none

CODEX_NETWORK_ACCESS

Allow network from sandbox

false

CODEX_WEB_SEARCH

disabled, cached, live

disabled

Instructions

Env Var

Description

CODEX_INSTRUCTIONS

Replaces the default instructions

CODEX_APPEND_INSTRUCTIONS

Appended to the default (usually what you want)

Advanced

Env Var

Description

CODEX_PERSIST_SESSION

true/false — enable session resume (default: true)

Authentication

Env Var

Description

Default

CODEX_API_KEY

OpenAI API key for this agent

inherited from parent

Security

  • Sandbox defaults to read-only — the agent can't write files unless you explicitly set workspace-write or danger-full-access.

  • cwd overrides preserve agent knowledge — when the host overrides cwd, the agent's configured base directory is automatically added to additionalDirectories.

  • Security overrides narrow, never widen — per-invocation sandboxMode and approvalPolicy can only tighten (e.g., workspace-writeread-only), never loosen.

  • _reply tool respects persistence — not registered when CODEX_PERSIST_SESSION=false.

  • API keys are redacted — the factory wizard never exposes CODEX_API_KEY in generated configs.

Architecture

┌─────────────────────────────────┐
│  MCP Client                     │
│  (Claude Desktop, Cursor, etc.) │
│                                 │
│  Sees: code_reviewer,           │
│        test_writer, quick_qa    │
└──────────┬──────────────────────┘
           │ JSON-RPC / stdio
┌──────────▼──────────────────────┐
│  Codex Octopus (per instance)   │
│                                 │
│  Env: CODEX_MODEL=o3            │
│       CODEX_SANDBOX_MODE=...    │
│       CODEX_APPEND_INSTRUCTIONS │
│                                 │
│  Calls: Codex SDK thread.run()  │
└──────────┬──────────────────────┘
           │ in-process
┌──────────▼──────────────────────┐
│  Codex SDK → Codex CLI          │
│  Runs autonomously: reads files,│
│  writes code, runs commands     │
│  Returns result + thread_id     │
└─────────────────────────────────┘

Known Limitations

  • minimal effort + web_search: OpenAI does not allow web_search tools with minimal reasoning effort. Use low or higher if web search is needed.

Development

pnpm install
pnpm build       # compile TypeScript
pnpm test        # run tests (vitest)
pnpm test:coverage  # coverage report

License

ISC - Xiaolai Li

Available Tools

2 tools
codexA

Send a task to an autonomous Codex agent. It reads/writes files, runs shell commands, searches codebases, and handles complex software engineering tasks end-to-end. Returns the result text plus a thread_id for follow-ups via codex_reply.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesTask or question for Codex
cwdNoWorking directory (overrides CODEX_CWD)
modelNoModel override (e.g. "gpt-5-codex", "o3", "codex-1")
additionalDirsNoExtra directories the agent can access for this invocation
effortNoReasoning effort override
sandboxModeNoSandbox mode override (can only tighten, never loosen)
approvalPolicyNoApproval policy override (can only tighten, never loosen)
networkAccessNoEnable network access from sandbox
webSearchModeNoWeb search mode
instructionsNoAdditional instructions (prepended to prompt)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Strong given no annotations: Explicitly discloses destructive capabilities ('reads/writes files', 'runs shell commands') and return format ('result text plus thread_id'). Would benefit from mention of execution duration, error handling, or sandbox safety implications beyond the parameter hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Perfect: Three sentences with zero waste. Front-loaded with core action ('Send a task'), middle sentence lists capabilities, final sentence covers return value and sibling relationship. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for complexity: With 10 parameters and 100% schema coverage, the description appropriately compensates for missing output schema by documenting return values (thread_id, result text). Mentions sibling relationship which is crucial. Minor gap: could clarify synchronous vs asynchronous behavior or lifecycle for such an autonomous agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Baseline score: With 100% schema description coverage, the description doesn't need to repeat parameter details. It adds context about the 'prompt' (task/question nature) implicitly through the agent description, but doesn't elaborate on specific parameters like 'approvalPolicy' or 'sandboxMode' semantics beyond the schema enums.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Excellent: 'Send a task to an autonomous Codex agent' provides specific verb+resource. Capabilities list ('reads/writes files, runs shell commands...') clearly scopes functionality. Explicitly distinguishes from sibling 'codex_reply' by noting this tool initiates tasks while the sibling handles follow-ups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Good: Establishes workflow relationship with sibling by stating returns include 'thread_id for follow-ups via codex_reply'. However, lacks explicit guidance on when NOT to use this vs alternatives or prerequisites (e.g., when to use codex_reply directly instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_replyA

Continue a previous codex conversation by thread ID. Use this for follow-up questions, iterative refinement, or multi-step workflows that build on prior context.

ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idYesThread ID from a prior codex response
promptYesFollow-up instruction or question
cwdNoWorking directory override
modelNoModel override

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It successfully conveys stateful behavior ('build on prior context') but lacks disclosure of error handling (invalid thread_id), side effects, output format, or rate limit behaviors expected for a conversational tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. First sentence front-loads the core action; second provides usage context. Every word earns its place with no redundancy or tautology.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 100% schema coverage and no output schema, the description adequately covers the tool's purpose and differentiation from siblings. Minor gap: no mention of error scenarios (e.g., expired threads) or what constitutes a valid thread_id format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, establishing baseline 3. The description does not redundantly explain parameters already well-documented in schema (thread_id, prompt), nor does it mention optional parameters (cwd, model), but no additional semantic clarification is required given complete schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the action ('Continue'), resource ('codex conversation'), and mechanism ('by thread ID'), clearly distinguishing it from the sibling 'codex' tool which likely initiates new conversations. The scope is precise and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance ('Use this for follow-up questions, iterative refinement, or multi-step workflows') that implies continuation versus new conversation. Could be strengthened by explicitly contrasting when NOT to use it versus the sibling 'codex' tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv1.0.1
    • First observedcodex
    • First observedcodex_reply

TDQS

A4/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: `codex` initiates a new autonomous task and returns a thread_id, while `codex_reply` explicitly requires that thread_id to continue an existing conversation. No functional overlap exists.

Naming Consistency4/5

Both tools share the `codex` prefix which groups them logically, and `codex_reply` clearly indicates its action. However, the base tool `codex` lacks a descriptive action verb (e.g., `codex_send` or `codex_start`) despite its description emphasizing sending tasks, creating a minor inconsistency in the naming pattern.

Tool Count3/5

With only 2 tools, the set is borderline thin for a server claiming to handle 'complex software engineering tasks end-to-end.' While the initiate-respond loop is functional, the lack of supporting operations (list threads, check status, cancel) makes the surface feel constrained rather than comprehensively scoped.

Completeness3/5

The core conversational workflow is covered (start thread, continue thread), but significant lifecycle operations are missing: there is no way to list active threads, retrieve thread history without continuing, cancel running tasks, or delete threads. These gaps limit agent ability to manage long-running or parallel autonomous tasks.

Maintenance

ActivityActive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    An MCP server that wraps OpenAI's Codex CLI to automate repository cloning and code analysis tasks. It enables users to execute complex coding requests on specific Git branches and subfolders using standardized MCP tools.
    11
    45
    MIT
  • F
    license
    B
    quality
    Not graded
    maintenance
    An MCP server for the OpenAI Codex CLI that provides coding assistance with multi-turn session management and reasoning depth control. It enables users to perform code analysis, generation, and refactoring through Claude with native resume support for conversational context.
    4
    765
    1
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/xiaolai/codex-octopus'

If you have feedback or need assistance with the MCP directory API, please join our Discord server