Codex Octopus
Wraps the OpenAI Codex SDK to provide specialized AI agents for coding tasks, allowing for the configuration of specific models, sandbox environments, and reasoning effort levels for use cases like code review and test generation.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Codex Octopuswrite thorough unit tests for the user authentication logic in auth.js"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Codex Octopus
One brain, many arms.
An MCP server that wraps the OpenAI Codex SDK, letting you run multiple specialized Codex agents — each with its own model, sandbox, effort, and personality — from any MCP client.
Why
Codex is powerful. But one instance does everything the same way. Sometimes you want a strict code reviewer in read-only sandbox. A test writer with workspace-write access. A cheap quick helper on minimal effort. A deep thinker on xhigh.
Codex Octopus lets you spin up as many of these as you need. Same binary, different configurations. Each one shows up as a separate tool in your MCP client.
Related MCP server: Codex MCP Server
Prerequisites
Node.js >= 18
Codex CLI — the Codex SDK spawns the Codex CLI under the hood, so you need it installed (
@openai/codex)OpenAI API key (
CODEX_API_KEYenv var) or inherited from parent process
Install
npm install codex-octopusOr use npx directly in your .mcp.json (see Quick Start below).
Quick Start
Add to your .mcp.json:
{
"mcpServers": {
"codex": {
"command": "npx",
"args": ["codex-octopus@latest"],
"env": {
"CODEX_SANDBOX_MODE": "workspace-write",
"CODEX_APPROVAL_POLICY": "never"
}
}
}
}This gives you two tools: codex and codex_reply. That's it — you have Codex as a tool.
Multiple Agents
The real power is running several instances with different configurations:
{
"mcpServers": {
"code-reviewer": {
"command": "npx",
"args": ["codex-octopus@latest"],
"env": {
"CODEX_TOOL_NAME": "code_reviewer",
"CODEX_SERVER_NAME": "code-reviewer",
"CODEX_DESCRIPTION": "Strict code reviewer. Read-only sandbox.",
"CODEX_MODEL": "o3",
"CODEX_SANDBOX_MODE": "read-only",
"CODEX_APPEND_INSTRUCTIONS": "You are a strict code reviewer. Report real bugs, not style preferences.",
"CODEX_EFFORT": "high"
}
},
"test-writer": {
"command": "npx",
"args": ["codex-octopus@latest"],
"env": {
"CODEX_TOOL_NAME": "test_writer",
"CODEX_SERVER_NAME": "test-writer",
"CODEX_DESCRIPTION": "Writes thorough tests with edge case coverage.",
"CODEX_MODEL": "gpt-5-codex",
"CODEX_SANDBOX_MODE": "workspace-write",
"CODEX_APPEND_INSTRUCTIONS": "Write tests first. Cover edge cases. TDD."
}
},
"quick-qa": {
"command": "npx",
"args": ["codex-octopus@latest"],
"env": {
"CODEX_TOOL_NAME": "quick_qa",
"CODEX_SERVER_NAME": "quick-qa",
"CODEX_DESCRIPTION": "Fast answers to quick coding questions.",
"CODEX_EFFORT": "minimal"
}
}
}
}Your MCP client now sees three distinct tools — code_reviewer, test_writer, quick_qa — each purpose-built.
Agent Factory
Don't want to write configs by hand? Add a factory instance:
{
"mcpServers": {
"agent-factory": {
"command": "npx",
"args": ["codex-octopus@latest"],
"env": {
"CODEX_FACTORY_ONLY": "true",
"CODEX_SERVER_NAME": "agent-factory"
}
}
}
}This exposes a single create_codex_mcp tool — an interactive wizard. Tell it what you want ("a strict code reviewer with read-only sandbox") and it generates the .mcp.json entry for you.
Tools
Each non-factory instance exposes:
Tool | Purpose |
| Send a task to the agent, get a response + |
| Continue a previous conversation by |
Per-invocation parameters (override server defaults):
Parameter | Description |
| The task or question (required) |
| Working directory override |
| Model override |
| Extra directories the agent can access |
| Reasoning effort ( |
| Sandbox override (can only tighten, never loosen) |
| Approval override (can only tighten, never loosen) |
| Enable network access from sandbox |
| Web search: |
| Additional instructions (prepended to prompt) |
Configuration
All configuration is via environment variables in .mcp.json. Every env var is optional.
Identity
Env Var | Description | Default |
| Tool name prefix ( |
|
| Tool description shown to the host AI | generic |
| MCP server name in protocol handshake |
|
| Only expose the factory wizard tool |
|
Agent
Env Var | Description | Default |
| Model ( | SDK default |
| Working directory |
|
|
|
|
|
|
|
|
| SDK default |
| Extra directories (comma-separated) | none |
| Allow network from sandbox |
|
|
|
|
Instructions
Env Var | Description |
| Replaces the default instructions |
| Appended to the default (usually what you want) |
Advanced
Env Var | Description |
|
|
Authentication
Env Var | Description | Default |
| OpenAI API key for this agent | inherited from parent |
Security
Sandbox defaults to
read-only— the agent can't write files unless you explicitly setworkspace-writeordanger-full-access.cwdoverrides preserve agent knowledge — when the host overridescwd, the agent's configured base directory is automatically added toadditionalDirectories.Security overrides narrow, never widen — per-invocation
sandboxModeandapprovalPolicycan only tighten (e.g.,workspace-write→read-only), never loosen._replytool respects persistence — not registered whenCODEX_PERSIST_SESSION=false.API keys are redacted — the factory wizard never exposes
CODEX_API_KEYin generated configs.
Architecture
┌─────────────────────────────────┐
│ MCP Client │
│ (Claude Desktop, Cursor, etc.) │
│ │
│ Sees: code_reviewer, │
│ test_writer, quick_qa │
└──────────┬──────────────────────┘
│ JSON-RPC / stdio
┌──────────▼──────────────────────┐
│ Codex Octopus (per instance) │
│ │
│ Env: CODEX_MODEL=o3 │
│ CODEX_SANDBOX_MODE=... │
│ CODEX_APPEND_INSTRUCTIONS │
│ │
│ Calls: Codex SDK thread.run() │
└──────────┬──────────────────────┘
│ in-process
┌──────────▼──────────────────────┐
│ Codex SDK → Codex CLI │
│ Runs autonomously: reads files,│
│ writes code, runs commands │
│ Returns result + thread_id │
└─────────────────────────────────┘Known Limitations
minimaleffort + web_search: OpenAI does not allowweb_searchtools withminimalreasoning effort. Uselowor higher if web search is needed.
Development
pnpm install
pnpm build # compile TypeScript
pnpm test # run tests (vitest)
pnpm test:coverage # coverage reportLicense
ISC - Xiaolai Li
Available Tools
2 toolscodexA
Send a task to an autonomous Codex agent. It reads/writes files, runs shell commands, searches codebases, and handles complex software engineering tasks end-to-end. Returns the result text plus a thread_id for follow-ups via codex_reply.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Task or question for Codex | |
| cwd | No | Working directory (overrides CODEX_CWD) | |
| model | No | Model override (e.g. "gpt-5-codex", "o3", "codex-1") | |
| additionalDirs | No | Extra directories the agent can access for this invocation | |
| effort | No | Reasoning effort override | |
| sandboxMode | No | Sandbox mode override (can only tighten, never loosen) | |
| approvalPolicy | No | Approval policy override (can only tighten, never loosen) | |
| networkAccess | No | Enable network access from sandbox | |
| webSearchMode | No | Web search mode | |
| instructions | No | Additional instructions (prepended to prompt) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Strong given no annotations: Explicitly discloses destructive capabilities ('reads/writes files', 'runs shell commands') and return format ('result text plus thread_id'). Would benefit from mention of execution duration, error handling, or sandbox safety implications beyond the parameter hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Perfect: Three sentences with zero waste. Front-loaded with core action ('Send a task'), middle sentence lists capabilities, final sentence covers return value and sibling relationship. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for complexity: With 10 parameters and 100% schema coverage, the description appropriately compensates for missing output schema by documenting return values (thread_id, result text). Mentions sibling relationship which is crucial. Minor gap: could clarify synchronous vs asynchronous behavior or lifecycle for such an autonomous agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Baseline score: With 100% schema description coverage, the description doesn't need to repeat parameter details. It adds context about the 'prompt' (task/question nature) implicitly through the agent description, but doesn't elaborate on specific parameters like 'approvalPolicy' or 'sandboxMode' semantics beyond the schema enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Excellent: 'Send a task to an autonomous Codex agent' provides specific verb+resource. Capabilities list ('reads/writes files, runs shell commands...') clearly scopes functionality. Explicitly distinguishes from sibling 'codex_reply' by noting this tool initiates tasks while the sibling handles follow-ups.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Good: Establishes workflow relationship with sibling by stating returns include 'thread_id for follow-ups via codex_reply'. However, lacks explicit guidance on when NOT to use this vs alternatives or prerequisites (e.g., when to use codex_reply directly instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_replyA
Continue a previous codex conversation by thread ID. Use this for follow-up questions, iterative refinement, or multi-step workflows that build on prior context.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | Yes | Thread ID from a prior codex response | |
| prompt | Yes | Follow-up instruction or question | |
| cwd | No | Working directory override | |
| model | No | Model override |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It successfully conveys stateful behavior ('build on prior context') but lacks disclosure of error handling (invalid thread_id), side effects, output format, or rate limit behaviors expected for a conversational tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. First sentence front-loads the core action; second provides usage context. Every word earns its place with no redundancy or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 100% schema coverage and no output schema, the description adequately covers the tool's purpose and differentiation from siblings. Minor gap: no mention of error scenarios (e.g., expired threads) or what constitutes a valid thread_id format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, establishing baseline 3. The description does not redundantly explain parameters already well-documented in schema (thread_id, prompt), nor does it mention optional parameters (cwd, model), but no additional semantic clarification is required given complete schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the action ('Continue'), resource ('codex conversation'), and mechanism ('by thread ID'), clearly distinguishing it from the sibling 'codex' tool which likely initiates new conversations. The scope is precise and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance ('Use this for follow-up questions, iterative refinement, or multi-step workflows') that implies continuation versus new conversation. Could be strengthened by explicitly contrasting when NOT to use it versus the sibling 'codex' tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.1- First observed
codex - First observed
codex_reply
TDQS
The two tools have clearly distinct purposes: `codex` initiates a new autonomous task and returns a thread_id, while `codex_reply` explicitly requires that thread_id to continue an existing conversation. No functional overlap exists.
Both tools share the `codex` prefix which groups them logically, and `codex_reply` clearly indicates its action. However, the base tool `codex` lacks a descriptive action verb (e.g., `codex_send` or `codex_start`) despite its description emphasizing sending tasks, creating a minor inconsistency in the naming pattern.
With only 2 tools, the set is borderline thin for a server claiming to handle 'complex software engineering tasks end-to-end.' While the initiate-respond loop is functional, the lack of supporting operations (list threads, check status, cancel) makes the surface feel constrained rather than comprehensively scoped.
The core conversational workflow is covered (start thread, continue thread), but significant lifecycle operations are missing: there is no way to list active threads, retrieve thread history without continuing, cancel running tasks, or delete threads. These gaps limit agent ability to manage long-running or parallel autonomous tasks.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for building and testing AI agents with multi-model experimentation and insights.
An MCP server that gives your AI access to the source code and docs of all public github repos
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
A MCP server built for developers enabling Git based project management with project and personal…
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceAn MCP server that wraps OpenAI's Codex CLI to automate repository cloning and code analysis tasks. It enables users to execute complex coding requests on specific Git branches and subfolders using standardized MCP tools.1145MIT
- FlicenseBqualityNot gradedmaintenanceAn MCP server for the OpenAI Codex CLI that provides coding assistance with multi-turn session management and reasoning depth control. It enables users to perform code analysis, generation, and refactoring through Claude with native resume support for conversational context.47651-
- AlicenseNot gradedqualityCmaintenanceA universal MCP server for spawning agents with any OpenAI-compatible LLM, supporting cloud and local models, and integrating with Claude Code, OpenCode, and Codex CLI.MIT
- AlicenseAqualityDmaintenanceAn MCP server that provides technical consultation, code review, and code explanation by integrating with OpenAI's Codex CLI, enabling AI-powered coding assistance in a sandboxed, read-only environment.3641MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/xiaolai/codex-octopus'
If you have feedback or need assistance with the MCP directory API, please join our Discord server