Skip to main content
Glama

Codex Supervisor MCP

Локальный мост на основе Model Context Protocol (MCP), позволяющий совместимому хосту запускать, отслеживать, направлять, прерывать, возобновлять и одобрять работу Codex app-server.

Он оборачивает codex app-server; он не автоматизирует терминальный интерфейс и не извлекает данные из IDE.

Возможности

Сервер предоставляет следующие MCP-инструменты:

Инструмент

Назначение

codex_start

Запустить новый поток Codex и ход в разрешённом репозитории.

codex_send

Отправить новую инструкцию, когда активный ход простаивает.

codex_steer

Добавить указания к активному ходу.

codex_status

Прочитать состояние потока, последние события, последнее сообщение агента, diff и ожидающие запросы.

codex_wait

Долгий опрос до завершения, ошибки, прерывания или запроса на одобрение.

codex_interrupt

Прервать активный ход.

codex_list_threads

Список сохранённых потоков внутри настроенных корней.

codex_read_thread

Прочитать сохранённый авторизованный поток.

codex_list_approvals

Просмотреть ожидающие запросы app-server.

codex_resolve_approval

Принять, отклонить или отменить одобрения на выполнение команд и изменение файлов.

Related MCP server: Codex Bridge MCP

Требования

  • Node.js 22 или новее.

  • Актуальный Codex CLI, доступный как codex.

  • Аутентификация Codex CLI уже настроена.

  • Один или несколько явных локальных корней репозиториев.

Этот проект не имеет npm-зависимостей времени выполнения.

Проверка пакета

npm test

Тестовый набор использует совместимый с протоколом имитатор app-server. Он не выполняет запросы к модели и не требует аутентификации Codex.

Установка с помощью Codex CLI

Используйте имя MCP-сервера codex-supervisor. Имя должно совпадать с CODEX_SUPERVISOR_MCP_NAME; мост использует его, чтобы дочерний app-server не загружал этот же MCP-сервер рекурсивно.

macOS или Linux

codex mcp add codex-supervisor \
  --env CODEX_ALLOWED_ROOTS="/Users/you/code:/Users/you/work" \
  --env CODEX_SUPERVISOR_MCP_NAME="codex-supervisor" \
  -- node "/absolute/path/to/codex-supervisor-mcp/src/index.mjs"

Используйте разделитель списка путей платформы между корнями. На macOS и Linux это двоеточие (:).

Windows PowerShell

codex mcp add codex-supervisor `
  --env CODEX_ALLOWED_ROOTS="C:\src;D:\work" `
  --env CODEX_SUPERVISOR_MCP_NAME="codex-supervisor" `
  -- node "C:\absolute\path\to\codex-supervisor-mcp\src\index.mjs"

В Windows разделитель списка путей — точка с запятой (;).

Подтвердите регистрацию:

codex mcp list

В Codex введите /mcp, чтобы проверить подключённый сервер.

Установка с помощью config.toml

Скопируйте и адаптируйте examples/config.toml, затем поместите его содержимое в ~/.codex/config.toml или в .codex/config.toml доверенного проекта.

Используйте абсолютные пути. Сохраняйте идентификатор сервера и CODEX_SUPERVISOR_MCP_NAME одинаковыми.

ChatGPT desktop или расширение Codex IDE

  1. Откройте Settings → MCP servers → Add server.

  2. Установите имя codex-supervisor.

  3. Выберите STDIO.

  4. Установите команду node.

  5. Добавьте абсолютный путь к src/index.mjs как единственный аргумент.

  6. Добавьте CODEX_ALLOWED_ROOTS и CODEX_SUPERVISOR_MCP_NAME=codex-supervisor.

  7. Сохраните и перезапустите хост.

  8. Введите /mcp, чтобы проверить инструменты.

Локальные STDIO MCP-серверы не загружаются обычными веб-чатами ChatGPT. Использование этого моста из веба требует отдельно развёрнутого, аутентифицированного удалённого MCP-сервиса или размещённого плагина.

Типичный рабочий процесс

Попросите MCP-хост:

Use codex_start in /absolute/path/to/repository to implement the requested
change. Use workspaceWrite, keep network access disabled, wait for progress,
show me every approval request before resolving it, and report the final diff
and test result.

Хост должен следовать этой последовательности:

codex_start -> codex_wait
  approval request -> inspect -> codex_resolve_approval -> codex_wait
  active correction -> codex_steer -> codex_wait
  completed -> codex_status
  later follow-up -> codex_send -> codex_wait

Каждый вызов start/send/steer/interrupt возвращает eventCursor. Передавайте его как afterSequence в codex_wait или codex_status, чтобы избежать повторного воспроизведения старых событий.

approvalPolicy принимает текущие wire-значения app-server: on-request (по умолчанию) и untrusted. Устаревшие значения onRequest и unlessTrusted принимаются мостом и нормализуются перед запросом к app-server.

Публичный API одобрений принимает decline, даже если релиз Codex app-server для этого запроса объявляет только cancel. В этом случае мост использует безопасный ответ отмены app-server и сообщает и запрошенное, и фактическое решения.

Конфигурация

Переменная

По умолчанию

Значение

CODEX_ALLOWED_ROOTS

Обязательно

Корни репозиториев, разделённые разделителем списка путей платформы.

CODEX_BIN

codex

Путь к нативному исполняемому файлу Codex. Заглушки Windows .cmd, .bat и .ps1 отклоняются.

CODEX_SUPERVISOR_MCP_NAME

codex-supervisor

Идентификатор конфигурации MCP, отключённый во вложенном app-server для предотвращения рекурсии.

CODEX_ALLOW_NETWORK

0

Установите 1, чтобы разрешить вызывающим запрашивать сетевой доступ.

CODEX_EVENT_LIMIT

1000

Количество событий в памяти, ограничено 100–10 000.

CODEX_SUPERVISOR_DEBUG

0

Установите 1, чтобы копировать stderr Codex app-server в stderr этого сервера.

CODEX_APP_SERVER_ARGS

Внутреннее безопасное значение по умолчанию

Расширенный JSON-массив, заменяющий все аргументы, передаваемые codex.

Аргументы app-server по умолчанию эквивалентны:

-c mcp_servers.<CODEX_SUPERVISOR_MCP_NAME>.enabled=false app-server

Переопределение CODEX_APP_SERVER_ARGS убирает эту защиту от рекурсии. Добавьте эквивалентное отключающее переопределение самостоятельно.

Модель безопасности

  • CODEX_ALLOWED_ROOTS обязателен.

  • Пути канонизируются с помощью realpath; выходы через симлинки отклоняются.

  • Codex получает ограниченный доступ на чтение к выбранному репозиторию и платформенным значениям по умолчанию.

  • workspaceWrite ограничивает корни для записи выбранным репозиторием.

  • dangerFullAccess не предоставляется.

  • Сетевой доступ требует и CODEX_ALLOW_NETWORK=1, и networkAccess: true в задаче.

  • У моста нет универсального, неизолированного shell-инструмента.

  • Одобрения на выполнение команд и изменение файлов должны разрешаться явно.

  • Потоки вне разрешённых корней отклоняются или фильтруются.

  • Полезные нагрузки событий ограничены по размеру перед хранением.

  • Сохранённые пути потоков повторно канонизируются при использовании; удалённые или заменённые пути репозиториев закрываются отказом.

  • Мутации в одном потоке, ответы на одобрения и повторные удалённые вызовы сериализуются или дедуплицируются, а не выполняются дважды.

  • Транспортные ошибки рекурсивно редактируются и ограничиваются по размеру перед пересечением границ STDIO или HTTP.

  • Учётные данные ретранслятора и удалённого сервера (BIOTELE_* и CODEX_REMOTE_*) удаляются из среды дочернего Codex.

  • Удалённые отправки результатов аутентифицируются через HMAC, кодируются base64url, разбиваются на ограниченные фрагменты и проверяются по длине и SHA-256 перед использованием. Кодирование защищает транспорт от контентных фильтров; это не шифрование.

Дочерний app-server по-прежнему наследует настройки процесса, не относящиеся к ретранслятору, и вашу более широкую конфигурацию Codex. Проверьте прочие секреты среды, приложения, навыки, хуки и настроенные MCP-серверы перед использованием с недоверенным кодом. Очистка окружения не является границей безопасности операционной системы: дочерний процесс, работающий от того же пользователя Windows, может намеренно запрашивать настройки уровня пользователя. Используйте отдельную учётную запись Windows, если такая угроза входит в область риска.

Поддерживаемые запросы одобрений

Этот релиз разрешает:

  • item/commandExecution/requestApproval

  • item/fileChange/requestApproval

Прочие запросы app-server остаются видимыми через codex_status и codex_list_approvals, но мост отказывается отвечать на них. Это предотвращает возможность молчаливого предоставления разрешений или передачи чувствительных пользовательских данных через универсальную конечную точку ответов.

Сохранение и мониторинг

Codex владеет сохранённой историей потоков. Мост хранит в памяти буферы потоковых событий, последние дельты и состояние ожидающих запросов. Перезапуск MCP-сервера очищает это временное состояние, но codex_list_threads и codex_read_thread могут восстановить авторизованные сохранённые потоки.

Разработка

npm test
node --check src/index.mjs

Структура проекта:

src/app-server-client.mjs  Codex app-server JSONL client
src/approval-policy.mjs    Approval-policy validation and legacy normalization
src/event-store.mjs        Bounded event, turn, and approval state
src/security.mjs           Repository-root policy
src/supervisor-service.mjs Codex lifecycle orchestration
src/tool-registry.mjs      MCP tool schemas and validation
src/mcp-server.mjs         Dual-era MCP STDIO transport
src/index.mjs              Entrypoint
test/                      Unit and integration tests

Лицензия

MIT

Совместимость с Codex App Server

Версия 1.0.3 удаляет устаревшие поля readOnly.access и workspaceWrite.readOnlyAccess из turn/start. Текущие релизы Codex App Server используют профили разрешений, когда клиенту нужны пользовательские ограниченные области чтения. Супервизор продолжает ограничивать корни для записи выбранным репозиторием и проверяет каждый каталог задачи на соответствие CODEX_ALLOWED_ROOTS.

Удалённый ретранслятор Hostinger

Версия 1.2.5 предоставляет совместимый с Hostinger ретранслятор для удалённого MCP-доступа ChatGPT:

ChatGPT -> OAuth bearer JWT -> Hostinger /mcp -> namespace-routed queue
  codex_*  -> outbound Windows local-agent -> Codex app-server
  reeves_* -> outbound Reeves Android agent -> accessibility service

Публичная конечная точка /mcp проверяет RS256 OAuth-токены доступа от внешнего поставщика идентификации. Агенты Windows и Reeves используют независимые HMAC- учётные данные только для исходящего опроса, статуса, получения аренды и отправки результатов. Ретранслятор Hostinger никогда не запускает Codex и не читает локальные репозитории.

Размещённый ретранслятор сохраняет все существующие инструменты codex_* и дополнительно предоставляет reeves_status, reeves_tap, reeves_swipe, reeves_type, reeves_back, reeves_home, reeves_recents, reeves_sequence и reeves_screenshot. Локальный реестр STDIO Codex остаётся только для Codex. Заявки агентов фильтруются по аутентифицированному идентификатору ключа; поля маршрутизации, предоставленные клиентом, игнорируются.

reeves_screenshot возвращает пиксели Android как стандартный MCP-блок содержимого image (image/png с данными base64) вместе с шириной, высотой, временем захвата, идентификатором агента и метаданными длины в байтах. Агент Android использует объявляемый ретранслятором протокол фрагментированных результатов, поэтому непригодный локальный путь Android не раскрывается, и каждый подписанный HTTP-запрос остаётся в пределах лимита тела ретранслятора.

reeves_sequence отправляет от 1 до 50 упорядоченных действий устройства в одной маршрутизированной задаче. Android выполняет локально действия tap, swipe, type, Back, Home, Recents, wait и screenshot, по умолчанию останавливается на первой ошибке и по умолчанию возвращает одно итоговое MCP- изображение. Результаты включают индексированные результаты действий и аддитивные, не содержащие секретов тайминги этапов ретранслятора/Android. Существующий 25-секундный запрос агента — это длинный опрос с пробуждением при постановке в очередь, а не задержка получения; Android сразу начинает новый запрос после каждой успешной отправки результата и повторно использует один пул соединений OkHttp.

Этот релиз также согласовывает поддерживаемую версию протокола MCP, выдаёт ограниченную сессию, привязанную к субъекту OAuth, и требует эту сессию в последующих запросах. Повторные вызовы инструментов привязываются к субъекту OAuth, MCP-сессии, типизированному JSON-RPC id и хэшу запроса; завершение сессии аннулирует её кэшированную или ожидающую работу. Релиз также очищает отменённую работу ретранслятора и состояние аварийно завершившегося app-server, повторно проверяет авторизованные пути, изолирует события по потокам и редактирует ограниченные вложенные данные об ошибках на каждом публичном транспорте.

Версия 1.2.5 также согласует codex_status.latestAgentMessage с авторизованной сохранённой стенограммой. Полностью сохранённые внешние завершения Codex, включая синтезированные ходы rollout-*, теперь заменяют устаревшие сообщения, наблюдаемые мостом, в то время как неполные или прерванные хвосты стенограммы остаются исключёнными.

Разверните обновлённый ретранслятор перед обновлением агента Windows. Новый ретранслятор по-прежнему принимает устаревшие одноразовые результаты, тогда как новый агент использует фрагментированный формат только после того, как ретранслятор объявит о его поддержке.

См. docs/REMOTE_DEPLOYMENT.md с шагами Hostinger hPanel, DNS для mcp.biotele.mx, настройкой Auth0, настройкой Microsoft Entra ID, настройкой и восстановлением веб-коннектора ChatGPT, переменными окружения, установкой локального агента и моделью угроз.

Available Tools

10 tools
codex_interruptInterrupt Codex turnA
DestructiveIdempotent

Request cancellation of an active Codex turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
turnIdNoOptional turn id; defaults to the active turn known by this bridge.
threadIdYesCodex thread id.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds the nuance that the cancellation applies to an 'active' turn, which implies it may not have effect on non-active turns. However, the annotations already declare destructiveHint and idempotentHint, so the description does not need to restate those. It does not clarify what happens if no active turn exists or whether this is a request versus a guaranteed cancellation, but with annotations covering the safety profile, the additional context is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that directly conveys the tool's purpose without any filler. It is perfectly concise and well-structured for such a simple operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the presence of annotations (destructive, idempotent) and full schema descriptions, the description provides sufficient context for invocation. It does not mention return values or error scenarios, but with no output schema and a straightforward operation, this is not a critical gap. It covers the essential action and target.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have thorough descriptions in the schema (100% coverage), including meaning and default behavior for turnId. The tool description itself adds no further parameter semantics, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action (request cancellation) on a specific resource (an active Codex turn), which distinguishes it from sibling tools like codex_start or codex_status. The verb 'request cancellation' is unambiguous and the phrase 'active Codex turn' scopes the operation appropriately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not mention that it should be used to stop an ongoing turn, or that one might use codex_status first to check for an active turn. The description simply states what it does without any contextual 'when-to-use' or exclusionary language.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_list_approvalsList Codex approvalsA
Read-onlyIdempotent

List pending app-server requests. This release resolves command-execution and file-change approval requests; unsupported request types remain visible.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdNoOptional authorized thread filter.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly and non-destructive behavior. The description adds context that unsupported request types remain visible and specifies which types this release resolves, going beyond annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states the primary action, the second adds scope clarification. No waste, information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with good annotations and full schema coverage, the description covers the operation and adds useful context about supported versus unsupported request types. It does not specify the return format, but no output schema exists, so the description is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, threadId, is fully documented in the schema as 'Optional authorized thread filter.' The description adds no further parameter semantics, so the schema carries the weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('pending app-server requests'), clearly distinguishing it from siblings like codex_resolve_approval and codex_list_threads. The title corroborates this.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly provides the context for when to use the tool (listing pending approvals). However, it does not explicitly mention alternatives or exclusions, so it does not achieve a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_list_threadsList Codex threadsA
Read-onlyIdempotent

List persisted Codex threads whose working directories are inside CODEX_ALLOWED_ROOTS. Results outside those roots are filtered.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional exact allowed working-directory filter.
limitNo
cursorNoOptional pagination cursor.
searchTermNoOptional thread search text.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and idempotent safety, so the description adds value by revealing the filtering behavior based on CODEX_ALLOWED_ROOTS and the 'persisted' qualifier. This goes beyond what annotations provide and gives important behavioral context for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences. The first sentence states the primary action and scope, and the second clarifies a key behavioral constraint. No unnecessary words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple listing tool with no output schema, the description covers the essential purpose and a key constraint. Combined with the schema and annotations, it provides sufficient context to use the tool, though it could mention pagination or return format explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not elaborate on any parameters. However, schema descriptions cover 3 of 4 parameters (cwd, cursor, searchTerm), and the remaining limit parameter has min/max constraints. With 75% schema coverage, the description's lack of parameter details is acceptable but not outstanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource ('persisted Codex threads') and a scoping constraint (working directories inside CODEX_ALLOWED_ROOTS). This clearly distinguishes it from sibling tools like codex_read_thread (which reads a single thread) and codex_list_approvals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating the tool lists threads and filters by allowed roots. It gives clear context (listing function) but does not explicitly name alternative tools or exclusions. The mention that results outside allowed roots are filtered provides actionable guidance about when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_read_threadRead Codex threadA
Read-onlyIdempotent

Read a persisted authorized Codex thread. The bridge denies threads outside CODEX_ALLOWED_ROOTS.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYesCodex thread id.
includeTurnsNoInclude full persisted turns. Omit for the safer compact thread view.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds a substantive behavioral detail: the bridge denies threads outside CODEX_ALLOWED_ROOTS, which is not visible in annotations. This informs the agent about potential authorization failures and scoping limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The main purpose is front-loaded in the first sentence, and the second sentence provides a crucial constraint. Every word contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, strong annotations, and full parameter schema coverage, the description is mostly complete. It does not describe the return value format, but with no output schema, that is a gap; however, the name and the includeTurns parameter reasonably convey what is returned. The authorization constraint adds essential context for operational use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the structured schema already defines threadId and includeTurns. The description adds no additional meaning beyond what the schema provides, such as the distinction between compact and full turn views, which is already in the schema. Baseline 3 is appropriate when schema covers parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Read') and the resource ('a persisted authorized Codex thread'), immediately distinguishing it from sibling tools like codex_list_threads (list threads) or codex_status (status check). The scope is well-defined with the authorization constraint, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for reading a specific thread's contents, but it does not explicitly differentiate from alternatives like codex_status or codex_list_threads. The mention of CODEX_ALLOWED_ROOTS is a constraint on valid inputs rather than guidance on when to choose this tool over others. Use is implied rather than clearly instructed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_resolve_approvalResolve Codex approvalA
Destructive

Explicitly accept, accept for the session, decline, or cancel a pending Codex command-execution or file-change approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
decisionYes
requestKeyYesOpaque requestKey returned by codex_status or codex_list_approvals.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark the tool as destructive (destructiveHint: true) and non-idempotent (idempotentHint: false). The description adds the specific decision actions (accept, acceptForSession, decline, cancel) but does not disclose behavioral details such as whether accepting triggers command execution or whether cancellation is reversible. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence that efficiently lists all decision actions and the approval types. There is no redundancy or filler, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool, the description covers the core purpose and decision spectrum adequately. However, it fails to mention how to obtain a requestKey or what happens after a decision (e.g., whether the command execution proceeds). The lack of an output schema means return values are unaddressed, but this is a minor gap given the simple nature of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is only 50%, with requestKey having a description and decision lacking one. The description repeats the enum values without explaining their semantic differences or the consequences of each choice. It adds no meaning beyond the schema for requestKey, merely referencing 'pending approval' without clarifying how to obtain or use the key.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses explicit verbs (accept, decline, cancel) and names the resource (pending Codex command-execution or file-change approval), making the tool's purpose immediately clear. It also distinguishes this from sibling tools like codex_list_approvals by indicating this tool resolves rather than lists approvals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when a pending approval exists, but it does not explicitly state when to use this tool versus alternatives, nor does it mention prerequisites like obtaining a requestKey. There is no guidance on when to choose acceptForSession over accept, or when to cancel rather than decline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_sendSend new Codex turnA
Destructive

Resume an authorized Codex thread when necessary and start a new turn. Use only when no turn is active; use codex_steer for an in-flight turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional allowed repository directory override.
modelNoOptional Codex model id. Omit to use the configured default.
effortNoOptional reasoning effort accepted by the selected model.
promptYesNew user instruction.
threadIdYesCodex thread id.
sandboxModeNoread-only prevents edits; workspace-write permits writes only in the repository. Legacy camelCase aliases remain accepted.workspace-write
networkAccessNoRequires CODEX_ALLOW_NETWORK=1 on the MCP server.
approvalPolicyNoControls when Codex asks before command execution or file changes.on-request

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructive and open-world behavior, so the description doesn't need to repeat those. However, it adds only a state constraint ('no turn active') without explaining what 'authorized' means, potential side effects like command execution, or how approval flow (given sibling codex_resolve_approval) interacts. Neutral but somewhat thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no fluff, front-loaded with action and usage. Every word earns its place and is highly scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While purpose and usage are clear, the tool has 8 parameters and no output schema. The description doesn't explain what the function returns, how to handle approval requests, or what distinguishes an 'authorized thread'. Sibling tools like codex_wait and codex_resolve_approval imply a larger workflow, but this description alone leaves some gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 8 parameters. The description adds no parameter-level detail beyond what the schema already provides, which is acceptable given the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Resume', 'start a new turn') with a clear resource ('Codex thread'), and differentiates itself from the sibling tool codex_steer by explicitly stating its scope. It clearly states what the tool does and when it applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('only when no turn is active') and names the alternative for the opposite case ('use codex_steer for an in-flight turn'). This is exactly the kind of direct usage guidance needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_startStart Codex taskA
Destructive

Start a new Codex app-server thread and immediately begin a turn in an allowed local repository. Return threadId, turnId, and an event cursor for codex_wait or codex_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYesExisting repository directory inside CODEX_ALLOWED_ROOTS.
modelNoOptional Codex model id. Omit to use the configured default.
effortNoOptional reasoning effort accepted by the selected model.
promptYesComplete implementation or investigation request for Codex.
sandboxModeNoread-only prevents edits; workspace-write permits writes only in the repository. Legacy camelCase aliases remain accepted.workspace-write
networkAccessNoRequires CODEX_ALLOW_NETWORK=1 on the MCP server.
approvalPolicyNoControls when Codex asks before command execution or file changes.on-request

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations: it returns threadId, turnId, and an event cursor, and clarifies that the turn begins immediately in an allowed local repository. Annotations already mark it as destructive and open-world, so no contradiction exists. It doesn't elaborate on approval or sandbox behavior, but those are covered by schema and annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences convey purpose, behavior, and return value without repetition. Every clause earns its place: it names the resource, the immediate action, the allowed scope, and the downstream tools that consume the return values.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description appropriately explains the return values. It also contextualizes the tool as a starting point for a thread lifecycle. It doesn't explicitly state that execution is asynchronous or that the returned IDs are needed for later calls, but those are strongly implied by the mention of the event cursor. The tool's complexity is well handled overall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents all parameters well. The description only adds 'allowed local repository' context that aligns with cwd, and mentions the return cursor, but doesn't add new details about parameters like model, effort, sandboxMode, etc. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Start a new Codex app-server thread'), the resource ('Codex app-server thread'), and the immediate behavior ('begin a turn in an allowed local repository'). It also distinguishes this from sibling tools by explicitly returning an event cursor for codex_wait or codex_status, signaling this is the entry point for new threads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what this tool does and implies its role relative to siblings by referencing post-start tools (codex_wait, codex_status). It doesn't explicitly say when not to use it or name alternatives for existing threads, but the context of 'new' thread is clear. The return cursor guidance helps the agent know how to follow up.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_statusRead Codex statusA
Read-onlyIdempotent

Read a thread snapshot, current turn state, pending approvals, latest agent message, latest diff, and recent streamed events.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYesCodex thread id.
maxEventsNo
includeTurnsNoInclude persisted turn history; this can produce a large response.
afterSequenceNoReturn events newer than this cursor.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds transparency by detailing the exact data returned (latest diff, streamed events, etc.) beyond the annotation flags. It does not mention potential large responses from includeTurns, but that is covered in the schema parameter description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the main action ('Read a thread snapshot') and then lists all data components in a logical sequence. No filler or redundant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately enumerates the key return categories (thread snapshot, turn state, approvals, etc.). It does not cover edge cases or polling behavior, but given the read-only annotation and schema parameters, it provides sufficient context for an agent to understand the tool's purpose and outputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not elaborate on any parameters, but the input schema already provides descriptions for 3 of 4 parameters (threadId, includeTurns, afterSequence). The only undocumented parameter, maxEvents, is self-explanatory. The description's phrase 'recent streamed events' loosely implies event-related parameters but adds minimal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a thread snapshot, current turn state, pending approvals, latest agent message, latest diff, and recent streamed events. It uses a specific verb ('Read') and resource, and the enumerated components distinguish it from sibling tools like codex_read_thread or codex_list_approvals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly conveys this is for getting a composite status view, but it does not explicitly state when to use this tool versus alternatives such as codex_read_thread or codex_list_approvals. No exclusions or alternative recommendations are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_steerSteer active Codex turnA
Destructive

Append guidance to the currently active turn without creating another turn. Supply expectedTurnId when available to prevent steering the wrong turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesAdditional in-flight guidance.
threadIdYesCodex thread id.
expectedTurnIdNoOptional active turn id returned by codex_start or codex_send.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint false, destructiveHint true, and idempotentHint false, so the agent knows this modifies state. The description adds context about the 'currently active turn' and the safety mechanism of expectedTurnId to avoid steering the wrong turn. Still, it does not explain failure modes or what happens if no active turn exists, so it provides only moderate additional transparency beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the core action, and every word earns its place. It avoids repetition and includes both the primary function and an important usage hint without excess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and the presence of detailed annotations, the description is adequate. It covers the key behavior and the critical safety parameter. A minor gap is that it does not explain return behavior or error conditions (e.g., what happens if no active turn exists), but since there is no output schema and the context is straightforward, this is not a significant omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning for expectedTurnId by explaining its purpose ('prevent steering the wrong turn'), which goes slightly beyond the schema. However, it offers no additional context for prompt or threadId beyond the existing schema descriptions. Overall, the added value is marginal, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool appends guidance to the currently active turn, using the specific verb 'append' and resource 'guidance to active turn'. It also distinguishes itself from creating another turn, differentiating it from siblings like codex_send. The mention of expectedTurnId to prevent steering the wrong turn further clarifies its unique role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you want to add guidance to an ongoing turn rather than starting a new one. The phrase 'without creating another turn' provides an exclusion, and the instruction to supply expectedTurnId when available gives practical guidance. However, it does not explicitly name alternative tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_waitWait for Codex progressA
Read-onlyIdempotent

Long-poll an active turn until it completes, fails, requests approval, is interrupted, or reaches the timeout. Continue with the returned eventCursor.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYesCodex thread id.
maxEventsNo
timeoutMsNo
afterSequenceNoCursor returned by the preceding Codex tool call.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so behavioral disclosure of safety is covered. The description adds valuable context about long-polling semantics, terminal states, timeout behavior, and the eventCursor continuation mechanism, which goes beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the action and outcome, and every phrase adds value. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple polling tool, the description covers purpose, terminal conditions, timeout, and continuation with eventCursor. There is no output schema, but the description gives enough hint about the return value. It could mention what events contain or how to handle approval requests, but those may be covered by sibling tools like codex_resolve_approval.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: threadId and afterSequence are described, while maxEvents and timeoutMs are not. The description mentions 'timeout' and 'eventCursor', providing partial meaning for timeoutMs and afterSequence, but does not fully compensate for the undocumented parameters. Parameter names are mostly self-explanatory, but a clearer mapping would help.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function with a specific verb ('Long-poll') and a specific resource ('an active turn'), listing all terminal states: completes, fails, requests approval, interrupted, or timeout. This distinguishes it from siblings like codex_status (which likely polls status) and codex_read_thread (which reads a thread).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: after an active turn exists, to wait for completion or interruption, and to continue using the returned eventCursor. It does not explicitly name alternatives or exclusions, but the context is clear enough that an agent can infer it should be used instead of codex_status for blocking waits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev1.2.3
    • Changedcodex_read_thread2 fields changed
      • changedInput schema / properties / includeTurns / default
        Previous value: -trueNew value: +false
      • addedInput schema / properties / includeTurns / description
        Added value: +"Include full persisted turns. Omit for the safer compact thread view."
  2. 10 tool updatesv1.0.3
    • First observedcodex_interrupt
    • First observedcodex_list_approvals
    • First observedcodex_list_threads
    • First observedcodex_read_thread
    • First observedcodex_resolve_approval
    • First observedcodex_send
    • First observedcodex_start
    • First observedcodex_status
    • First observedcodex_steer
    • First observedcodex_wait

TDQS

A3.9/5.0
Disambiguation4/5

Most tools are clearly distinct, with start/send/steer/wait/interrupt targeting different phases of thread execution. The only notable overlap is between codex_read_thread and codex_status, which both provide access to thread content, though status is a broader snapshot and read_thread is more focused.

Naming Consistency3/5

All tools share the codex_ prefix, but the pattern is inconsistent: some use verb_noun (read_thread, list_approvals, resolve_approval), while others are bare verbs (start, send, steer, wait, interrupt) or a noun (status). The naming is predictable but not uniformly structured.

Tool Count5/5

With 10 tools, the server is well-scoped for its purpose of supervising Codex threads and approvals. Each tool covers a necessary action without unnecessary redundancy, fitting comfortably within the ideal range.

Completeness5/5

The tool surface provides full lifecycle coverage: starting, resuming, steering, monitoring, interrupting, and listing threads, plus complete approval management (list and resolve). No obvious dead ends or missing operations for this domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/redmikarimo/codex-supervisor-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server