Skip to main content
Glama

claude-max-mcp

An MCP server that turns your Claude Max / Pro subscription into a backend any AI agent can call.

If you pay for Claude Max or Pro, you have a real allotment of Claude usage already — but most agent tools (Cline, Continue, Cursor, custom MCP setups, even Claude Code nested inside another agent) hit the Anthropic API directly with separate credit-card billing. You end up paying twice.

This is a small stdio MCP server that fixes that. It exposes a single tool — ask_claude — which spawns Claude Code under the hood and uses your existing Claude.ai OAuth session to route the call. Net effect: the call counts against your Max/Pro subscription credits, not against any API key.

  Your MCP-compatible agent
  (Cline • Continue • Cursor • Claude Code itself •
   Hermes Agent • custom agents • etc.)
              │
              │  ask_claude(prompt)
              ▼
        ┌──────────────────┐
        │  claude-max-mcp  │  ← this repo
        │  (stdio server)  │
        └──────────────────┘
              │
              │  spawn  claude -p --model claude-opus-4-7
              ▼
        ┌──────────────────┐
        │  Claude Code CLI │
        └──────────────────┘
              │
              │  OAuth via ~/.claude.json
              ▼
        Claude Max / Pro subscription
        (your existing monthly allotment)

Why this exists

Three real scenarios this solves:

  1. Reduce duplicate billing. Stop paying for Max and API credits when the agent work could just count against Max.

  2. Prompt-enhancement router → Claude delegate. A cheap router model (Gemini, Sonnet, Haiku) does no substantive reasoning — it just adds context (project name, today's date, scope hints) to the user's message and forwards it verbatim to Claude Opus 4.7. Claude does all the actual thinking, billed to your Max subscription. Reference template in examples/hermes/SOUL-additions.md.

  3. Nest agents inside Claude Code. Use this from within a Claude Code session so a sub-agent can make its own scoped Claude call without polluting the parent session's context.

Related MCP server: OpenAI Agents MCP Server

Install

# 1. Install Claude Code CLI if you don't have it
npm install -g @anthropic-ai/claude-code

# 2. Log in once (OAuth — uses your Claude Max / Pro subscription)
claude /login

# 3. Install this MCP server
npm install -g claude-max-mcp
# or run from source:
#   git clone https://github.com/Jonahbkerr/claude-max-mcp
#   cd claude-max-mcp && npm install

# 4. Smoke test
claude-max-mcp --smoke
# → ✓ claude-max-mcp smoke test passed (model: claude-opus-4-7)

Register with your MCP client

The server is stdio. Drop it into your client's MCP config.

Claude Code itself

Add to ~/.claude/settings.json under mcpServers:

{
  "mcpServers": {
    "claude-max": { "command": "npx", "args": ["-y", "claude-max-mcp"] }
  }
}

Cline (VS Code)

In Cline's MCP server settings:

{
  "mcpServers": {
    "claude-max": {
      "command": "npx",
      "args": ["-y", "claude-max-mcp"]
    }
  }
}

Continue (VS Code / JetBrains)

In ~/.continue/config.yaml:

mcpServers:
  - name: claude-max
    command: npx
    args: [-y, claude-max-mcp]

Any other MCP-compatible client

Same pattern — command: npx, args: [-y, claude-max-mcp]. See examples/ for ready-to-paste configs.

Configure

Set on the MCP client side when you register the server:

Env var

Default

What it does

CLAUDE_MAX_MCP_MODEL

claude-opus-4-7

Model passed to claude --model. Aliases work (opus, sonnet, haiku).

CLAUDE_MAX_MCP_BIN

claude (from PATH)

Path to the Claude Code CLI.

CLAUDE_MAX_MCP_TIMEOUT_MS

600000 (10 min)

How long a single delegated call can run.

CLAUDE_MAX_MCP_CWD

$HOME

Working dir for the claude subprocess — point at a project root if you want Claude to see CLAUDE.md / files there.

CLAUDE_MAX_MCP_TOOL_NAME

ask_claude

Override the exposed tool name (useful when running multiple instances).

CLAUDE_MAX_MCP_EXTRA_ARGS

(empty)

Extra space-separated args to pass to claude -p.

You can register the same server multiple times under different names with different env to expose multiple tools to the same agent (e.g. one Opus instance scoped to a project, one Haiku for one-liners). See examples/hermes/multi-instance.sh.

How it works

docs/architecture.md — data flow, OAuth discovery, model selection, why subprocess.

docs/faq.md — Max sub credit math, common errors, troubleshooting, what this does and doesn't do.

Honest tradeoffs

  • Not free. Each call debits your Claude subscription allotment. The savings come from avoiding additional API billing, not from making Claude free.

  • Opus is ~5× Sonnet in credit terms. Burning thousands of Opus delegations per day will exhaust Max usage caps. Use CLAUDE_MAX_MCP_MODEL=sonnet or =haiku for routine work.

  • Subprocess overhead ≈ 1–2s per call. Fine for chat / agent latency, not for sub-second tool loops.

  • Daemon-managed clients (launchd, systemd) often strip PATH — set PATH or CLAUDE_MAX_MCP_BIN explicitly in env config.

Security

  • Anthropic credentials never pass through this server. It shells out to claude, which reads its own OAuth from ~/.claude.json.

  • The MCP surface is one tool, one string argument. No file I/O, no shell injection (args go through spawn, not a shell), no recursive tool-use loops.

  • To restrict what Claude can do during delegated calls, set CLAUDE_MAX_MCP_CWD to a sandboxed directory and use CLAUDE_MAX_MCP_EXTRA_ARGS="--allowed-tools=Read,Glob,Grep" (or similar) to scope its toolset.

License

MIT — see LICENSE.

Issues, PRs, and "this should also work with X" reports welcome.

Available Tools

1 tool
ask_claudeA

Delegate a substantive task to Claude Code, billed against the user's Claude Max / Pro subscription (model: claude-opus-4-7). Use this for: code generation, deep reasoning, multi-step engineering, refactoring, or anything that needs more capability than the routing model can reliably deliver. The receiver has NO memory of the calling agent's conversation — include everything Claude needs in the prompt argument: the user's verbatim request, project / file context, relevant prior messages, and what kind of output you want back. Returns Claude's full text response.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe full prompt to send to Claude. Claude has no access to the calling agent's conversation — include all context here.

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It discloses key behaviors: no memory of conversation, billing against subscription, and the model used. It lacks details on potential rate limits or response format beyond 'full text response', but overall is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two sentences—with no wasted words. It front-loads the core purpose and then provides essential usage guidance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has only one parameter and no output schema, the description is fully complete. It explains the tool's purpose, when to use it, how to structure the input, and the key behavioral constraint (no memory). No additional information is necessary for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter. The description adds significant meaning: it explains why the prompt must contain all context (no conversation access) and what elements it should include, going well beyond the schema's basic description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: delegating tasks to Claude Code. It specifies the model (claude-opus-4-7) and the subscription billing context, and contrasts with the routing model, providing strong clarity even without sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly lists when to use the tool (code generation, deep reasoning, multi-step engineering, refactoring) and what to include in the prompt (user request, context, prior messages, desired output). It also warns that Claude has no conversation memory, guiding correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev0.1.1
    • First observedask_claude

TDQS

A4.8/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap. The tool's purpose is distinct and unambiguous.

Naming Consistency5/5

The single tool name 'ask_claude' follows a clear verb_noun pattern. Consistency is perfect due to the lack of other tools.

Tool Count4/5

The server contains only one tool, which is appropriate for its narrow scope of delegating tasks to Claude Code. While a slightly larger set might be expected for a full SDK, the single tool is well-justified and not trivial.

Completeness5/5

The tool fully covers the stated domain of delegating substantive tasks to Claude Code. No additional tools are necessary for the server's purpose.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Jonahbkerr/claude-max-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server