ultimate-debate-mcp
Provides adversarial collaboration with GPT-5.2, enabling automated critique, verification, and debate of code, plans, or text.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ultimate-debate-mcpVerify the SQL injection fix is complete"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ultimate Debate MCP Server
A Model Context Protocol (MCP) server that enables Claude + GPT adversarial collaboration. GPT-5.2 automatically critiques, debates, and verifies Claude's work for higher-quality outputs.
Features
gpt_critique- Single-pass adversarial critique of code, plans, or textgpt_verify- Verify that fixes actually resolve identified issuesdebate- Multi-round debate loop between Claude and GPT
Related MCP server: Claude-Gemini MCP Integration Server
Requirements
Node.js >= 20.0.0
OpenAI API key with GPT-5.2 access
Claude Code CLI
Installation
1. Clone and Build
git clone https://github.com/YOUR_USERNAME/ultimate-debate-mcp.git
cd ultimate-debate-mcp
npm install
npm run build2. Install Globally in Claude Code
# Add the MCP server globally (works in any folder)
claude mcp add ultimate-debate --scope user -- node /path/to/ultimate-debate-mcp/dist/index.js3. Configure Environment Variables
Edit your ~/.claude.json and find the mcpServers section at the root level. Add your OpenAI API key:
{
"mcpServers": {
"ultimate-debate": {
"type": "stdio",
"command": "node",
"args": ["/path/to/ultimate-debate-mcp/dist/index.js"],
"env": {
"OPENAI_API_KEY": "sk-your-openai-api-key",
"OPENAI_MODEL": "gpt-5.2",
"RATE_LIMIT_RPM": "100",
"RATE_LIMIT_RPH": "1000",
"CACHE_MAX_SIZE": "100",
"CACHE_TTL_MS": "3600000"
}
}
}
}4. Restart Claude Code
Close and reopen Claude Code. Verify installation:
/mcpYou should see ultimate-debate listed with 3 tools.
Configuration Options
Variable | Default | Description |
| required | Your OpenAI API key |
|
| Model to use (gpt-5.2, gpt-4o, etc.) |
|
| Max requests per minute |
|
| Max requests per hour |
|
| Max cached responses |
|
| Cache TTL (1 hour) |
Auto-Collaboration Setup
To make Claude automatically use GPT review, add these rules to your global ~/.claude/CLAUDE.md:
# Auto-Collaboration Rules
When the `ultimate-debate` MCP server is available, automatically:
1. **Before implementation** (>20 lines): Call `gpt_critique` on the plan
2. **After implementation**: Call `gpt_critique` on the code
3. **After fixing issues**: Call `gpt_verify` to confirm resolution
4. **For architecture decisions**: Call `debate` with 2 rounds
## Decision Criteria
| Change Type | Action |
|-------------|--------|
| New feature | `debate` (2 rounds) |
| Bug fix | `gpt_critique` + `gpt_verify` |
| Refactor | `gpt_critique` |
| Trivial (<10 lines) | Skip review |
## Skip Review
Say "skip review" or "no gpt" to bypass auto-collaboration.Usage Examples
Manual Critique
Ask Claude to review code:
Critique this authentication code for security issuesManual Debate
For architecture decisions:
Debate whether we should use Redis or PostgreSQL for session storageManual Verify
After fixing issues:
Verify the SQL injection fix is completeMCP Tools Reference
gpt_critique
Single-pass adversarial critique.
Parameters:
content(required) - The content to critiquecontext(optional) - Additional contextfocus_areas(optional) - Array of:logic,security,performance,edge_cases,assumptions,completeness
Returns: Issues with severity (critical/high/medium/low), category, and suggestions.
gpt_verify
Verify fixes are resolved.
Parameters:
original(required) - Original content before fixrevised(required) - Revised content after fixissues_to_check(required) - Array of issues to verify
Returns: Lists of verified, unresolved, and new issues.
debate
Multi-round debate loop.
Parameters:
goal(required) - The objective to debatecontext(optional) - Additional contextartifacts(optional) - Code, diffs, specs, logsrounds(default: 2) - Number of debate rounds (1-5)confidence_threshold(default: 0.8) - Stop early if confidence reachedbudget(optional) - Token and cost limits
Returns: Critical issues, resolved issues, suggested patches, test recommendations.
Development
npm install # Install dependencies
npm run build # Build for production
npm run dev # Development mode with hot reload
npm run test # Run tests
npm run lint # Run linterLicense
MIT
Available Tools
3 toolsdebateA
Runs a full multi-round debate loop where ChatGPT critiques content and suggests revisions. Returns structured issues, patches, and test recommendations.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | The objective or question to debate | |
| budget | No | Token and cost budget | |
| rounds | No | Number of debate rounds (1-5, default 2) | |
| context | No | Additional context for the debate | |
| artifacts | No | Code, diffs, specs, or logs to analyze | |
| confidence_threshold | No | Confidence level to stop early (0-1, default 0.8) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose the key process: it runs a multi-round debate loop, critiques content, suggests revisions, and returns structured issues, patches, and test recommendations. It does not mention cost/token implications, early stopping via confidence_threshold, or exact internal iteration semantics, but the core behavioral pattern is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. It front-loads the main action and scope, then states the output format, making it easy to parse quickly without losing critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with nested objects and no output schema, the description communicates the basic behavior and return categories but not how budget, rounds, artifacts, or confidence_threshold interact with the debate loop. It also omits explicit usage guidance and output structure detail, though the input schema covers parameter meaning well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including goal, budget, rounds, context, artifacts, and confidence_threshold already has a description in the schema. The tool description itself adds no extra parameter-level meaning, which is acceptable under the baseline because the schema fully covers the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Runs a full multi-round debate loop') and a specific resource ('ChatGPT critiques content and suggests revisions'), making the tool's function immediately clear. It also distinguishes itself from the sibling tools gpt_critique and gpt_verify by emphasizing the multi-round, iterative nature rather than a single critique or verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by the phrase 'full multi-round debate loop,' which suggests this tool is for iterative critique-and-revise sessions rather than one-off critique or verification. However, there is no explicit when-to-use guidance and no mention of alternative tools or exclusion criteria, so the agent must infer when to choose this over gpt_critique or gpt_verify.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gpt_critiqueB
Single-pass adversarial critique of content using ChatGPT. Identifies logical gaps, edge cases, security issues, and incorrect assumptions.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The content to critique (code, text, plan) | |
| context | No | Additional context for better critique | |
| focus_areas | No | Specific areas to focus critique on |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds useful behavioral context such as 'single-pass' and 'adversarial', but it does not disclose the output format, limitations, or whether the input is mutated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences that front-load the core purpose and immediately enumerate the kinds of issues identified. There is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with well-described parameters, the description conveys the core action and output types. However, without an output schema, it does not clarify what the critique response looks like, and it misses any mention of the sibling tools or alternative use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides descriptions for all three parameters (content, context, focus_areas), yielding roughly 100% coverage, so the baseline is 3. The description's mention of issue types loosely maps to the focus_areas enum but adds no additional parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('critique') and resource ('content'), and enumerates specific output types (logical gaps, edge cases, security issues, incorrect assumptions). However, it does not explicitly distinguish the tool from sibling tools gpt_verify or debate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus its siblings, nor does it state when not to use it or any prerequisites. The intended use is only implied by the word 'critique'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gpt_verifyA
Strict verification that previously identified issues have been resolved in revised content. Also detects any new issues introduced.
| Name | Required | Description | Default |
|---|---|---|---|
| revised | Yes | Revised content to verify | |
| original | Yes | Original content before revision | |
| issues_to_check | Yes | Specific issues to verify are resolved |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the core behaviors (verifying resolution and detecting new issues) and the strictness of verification, but does not explain output format, how unresolved issues are reported, or what 'strict' means operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary purpose is front-loaded, and the secondary behavior is a single clarifying clause. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description should cover what the agent can expect in return. It doesn't state whether the tool returns a verdict, a list of remaining issues, severity levels, or structured data. Inputs are clear, but the missing output expectations create a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for all three parameters, and their descriptions already define original, revised, and issues_to_check. The description adds no extra parameter-level nuance beyond framing the overall task; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies that previously identified issues are resolved in revised content, and adds a second distinct behavior of detecting new issues. This is a specific verb+resource combination that inherently distinguishes it from siblings gpt_critique and debate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it when you have original content, revised content, and a specific list of previously identified issues to verify. It doesn't explicitly name alternatives or state conditions when not to use it, but the usage context is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v1.0.0- First observed
debate - First observed
gpt_critique - First observed
gpt_verify
TDQS
gpt_critique and debate both involve critique, so there is some boundary overlap, but the descriptions clearly separate single-pass critique from a full multi-round debate loop. gpt_verify is distinctly focused on checking whether prior issues were resolved, making misselection unlikely.
Two tools follow a gpt_<verb> pattern (gpt_critique, gpt_verify), while debate breaks the prefix pattern but still uses a single clear lowercase verb/noun. The naming is generally consistent and readable, with only the missing prefix creating minor inconsistency.
Three tools is a tight, well-scoped set for a debate-focused server. Each tool covers a distinct mode—single critique, verification of revisions, and full multi-round debate—so none feel redundant or unnecessary.
The surface covers the core debate workflow: critique content, run a full debate loop with revisions and patch suggestions, and verify that issues are resolved. There are no obvious dead ends for an agent trying to critique, revise, and validate content.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
AI code review for GitHub PRs with an MCP autofix loop for Claude Code and Cursor
Multi-LLM council: 25+ frontier models in parallel, consensus scoring, verdict-first code review.
- AxisOAuthdev.useaxis
Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.
Multi-model AI debates: GPT-4o, Claude, Gemini & 200+ models discuss, then synthesize insight.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables Claude to collaborate with Gemini for code reviews, second opinions, and iterative software development. It facilitates multi-step workflows including PRD creation and code generation through an AI orchestration framework.2181MIT
- AlicenseNot gradedqualityDmaintenanceBridges Claude Code and Google's Gemini AI models to enable AI-to-AI collaboration for code reviews, brainstorming, and direct questions.5MIT
- AlicenseNot gradedqualityCmaintenanceEnables running position-driven adversarial debates and code reviews between AI agents via MCP tools, supporting custom positions, multiple rounds, and local CLI models like Claude, Codex, and Gemini.581MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code to debate with Gemini on a given topic by managing conversation history and assigning personas, allowing users to watch AI agents argue.-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/captdevc/ultimate-debate-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server