Skip to main content
Glama

Thinking Agent MCP

License: MIT Node.js TypeScript MCP

Extend your thinking model's chain of thought via MCP tools — a Model Context Protocol server that exposes chat_agent, create_branch, and get_branch_details tools, enabling thinking models to offload subtasks to non-thinking models and build tree-structured multi-perspective analysis.


Features

  • 🧠 Chain of Thought Extension — Thinking models can delegate reasoning subtasks to non-thinking models via chat_agent, extending effective reasoning depth beyond single-model token limits

  • 🌳 Tree-Structured Thinkingcreate_branch enables recursive, multi-perspective exploration with four branch types (drill down / verify / explore / stash)

  • 🔍 Full Traceabilityget_branch_details retrieves the complete raw reasoning process of any created branch

  • 🛡️ Context Isolation — Tools are stateless and self-contained; all context must be packed into input_text. No conversation history dependency

  • 🎛️ Parameter Control — Fine-grained control over tool model output via temperature, top_p, seed, stop, and max_tokens

  • 🔌 Dual API Support — Works with both DeepSeek official API (recommended) and SiliconFlow API


Related MCP server: Long Reasoning MCP Server

Table of Contents


Quick Start

# Clone and install
git clone https://github.com/ScarletLilith/DeepSeekV4Flash_Thinking_TreeMCP.git
cd meditatorMCP
npm install

# Configure API (see Configuration section below)
# edit test/config.json or set environment variables

# Start the server
npm run build
npm start

# Or run development mode
npm run dev

Configuration

Configuration is loaded with the following priority: Environment variables > test/config.json

export DEEPSEEK_API_KEY=sk-your-key
export DEEPSEEK_BASE_URL=https://api.deepseek.com
export DEEPSEEK_MODEL=deepseek-v4-pro

Note: DeepSeek's thinking mode uses thinking: {type: "enabled"} (not enable_thinking: true).

Option 2: SiliconFlow API (Fallback)

export SILICONFLOW_API_KEY=sk-your-key
export SILICONFLOW_BASE_URL=https://api.siliconflow.cn/v1
export SILICONFLOW_MODEL=deepseek-ai/DeepSeek-V4-Flash

Config File

Create test/config.json (gitignored automatically):

{
  "baseUrl": "https://api.deepseek.com",
  "model": "deepseek-v4-pro",
  "apiKey": "sk-xxx"
}

Tools

chat_agent

Calls a non-thinking model to execute an independent subtask, extending the thinking model's chain of thought.

Parameters

Parameter

Type

Default

Description

input_text

string

required

Complete, self-contained task description with all context

system_prompt

string

optional

System prompt for role/behavior constraints

temperature

number

0.7

Sampling temperature (0.0–2.0). Low = precise, high = creative

top_p

number

0.9

Nucleus sampling threshold (0.0–1.0)

max_tokens

number

4096

Maximum output tokens (enforced server-side via API)

stop

string[]

[]

Stop sequences; empty array = natural completion

seed

number

optional

Random seed for reproducible output (with low temperature)

Parameter Strategies

Verification:  temperature=0.1, top_p=0.1,  max_tokens=2048, seed=42
Exploration:   temperature=1.2, top_p=0.95, max_tokens=4096
Balanced:      temperature=0.5, top_p=0.8,  max_tokens=4096

create_branch

Creates a thinking branch node with recursive nesting support for deep multi-perspective analysis.

Parameters

Parameter

Type

Default

Description

session_id

string

required

Session ID, consistent within a single reasoning session

input_text

string

required

Self-contained subtask description (≥30 characters)

call_type

string

drill_down

Branch type: drill_down / verify / explore / stash

parent_node_id

string

trunk

Parent node ID for tree nesting

Four Branch Types:

Type

Temperature

Purpose

drill_down

0.2

Deep-dive into a subproblem with focused precision

verify

0.0

Verify a conclusion or hypothesis with maximal determinism

explore

1.0

Divergent thinking from different angles with high creativity

stash

0.6

Temporarily record intermediate thoughts for later reference

Response

{
  "status": "success",
  "node_id": "n_a1b2c3d4",
  "conclusion": "The extracted conclusion text...",
  "confidence": 0.85,
  "remaining_quota": 12,
  "suggestions": [
    "发散探索完成,可对有价值的方向用 drill_down 深入",
    "还可创建 12 个分支,建议继续多角度探索"
  ]
}

get_branch_details

Retrieves the complete raw reasoning process of a previously created branch node.

Parameters

Parameter

Type

Default

Description

session_id

string

required

Session ID

node_id

string

required

Branch node ID returned by create_branch

Response

{
  "status": "success",
  "node_id": "n_a1b2c3d4",
  "raw_process": "The complete raw reasoning output from the model..."
}

Error Handling

Tools return structured errors with type and action fields for the thinking model to make informed decisions:

{
  "success": false,
  "type": "api",
  "action": "report",
  "error": "API authentication failed (401)",
  "status_code": 401
}

Error Type

action

Trigger

network

retry

DNS resolution failure, connection refused

api

backoff

429 rate limited

api

report

401 authentication failure

api

retry

5xx server errors

validation

fix_input

Empty input_text

config

report

Missing API Key / Model configuration

The server-side retry mechanism uses exponential backoff with jitter (1s→3s→7s, max 3 retries) for 429 and 5xx errors. Network errors (ENOTFOUND, ECONNREFUSED, ECONNRESET) are not automatically retried.


MCP Client Setup

Claude Desktop

{
  "mcpServers": {
    "thinking-agent": {
      "command": "node",
      "args": ["path/to/meditatorMCP/dist/index.js"],
      "env": {
        "DEEPSEEK_API_KEY": "sk-your-key",
        "DEEPSEEK_BASE_URL": "https://api.deepseek.com",
        "DEEPSEEK_MODEL": "deepseek-v4-pro"
      }
    }
  }
}

Any MCP-compatible Client

Configure stdio transport to point to node dist/index.js in the project directory, with the required environment variables set.


Testing

The project includes both interactive and automated test frameworks:

# Interactive CLI (with tools mode)
npm run test:with-tool

# Interactive CLI (pure thinking, no tools)
npm run test:without-tool

# Automated comparison test (runs both scenarios + generates report)
npm run test:comparison

# Batch end-to-end tests
npm run test:batch

Test Scripts

Script

Description

test/testFramework.ts

Interactive CLI test framework

test/comparisonTest.ts

Automated A/B comparison (with-tool vs without-tool)

test/runA.js

Scenario A: thinking model + tools (standalone, DeepSeek)

test/runB.js

Scenario B: pure thinking model (standalone, DeepSeek)

test/batchTest.ts

Batch end-to-end tests

Scoring: Each question is evaluated against 10 objective checkpoints (50 total). Evaluation is done by human reviewers, not automated scripts.

Note: The test/config.json file contains your API key and is automatically gitignored.


Project Structure

├── src/
│   ├── index.ts           # MCP Server entry point
│   ├── chatAgentTool.ts   # Tool implementations (chat_agent, create_branch, get_branch_details)
│   ├── gatekeeper.ts      # Input validation and quota enforcement
│   ├── strategyEngine.ts  # Parameter strategy mapping (call_type → temperature/top_p)
│   ├── nodeStore.ts       # Branch node storage and conclusion extraction
│   ├── schemas.ts         # Zod validation schemas and TypeScript types
│   ├── logger.ts          # Structured logging to stderr
│   └── polyfill.ts        # Node 14 fetch polyfill
├── test/
│   ├── comparisonTest.ts  # A/B comparison test
│   ├── testFramework.ts   # Interactive CLI test framework
│   ├── batchTest.ts       # Batch testing
│   ├── runA.js            # Scenario A test (DeepSeek)
│   ├── runB.js            # Scenario B test (DeepSeek)
│   └── config.json        # API configuration (gitignored)
├── .env.example           # Environment variable template
├── blueprint.md           # Project design blueprint (Chinese)
├── package.json
├── tsconfig.json
└── README.md

Development

# Build TypeScript
npm run build

# Start production server
npm run start

# Development mode (ts-node, no build step)
npm run dev

Design Philosophy

  1. Self-Contained Task Descriptions — All context must be packed into input_text; tools never rely on conversation history

  2. Context Isolation — Each tool call is stateless and independent, preventing context explosion in the main chain

  3. Token Cost Optimization — Context is consumed by the cheaper non-thinking model's input tokens, not the thinking model's output tokens

  4. Tree-Structured Reasoning — Complex problems are decomposed into independent branches, each analyzed separately, then synthesized


Benchmark: MCP Tools Impact on Output Quality

We conducted a controlled experiment comparing 3 approaches across 5 challenging engineering problems (distributed consensus, service mesh, RTOS kernel, columnar storage engine, multi-modal AI agent framework).

Test Groups

Group

Model

API

Tools

A

GLM-5.2

SiliconFlow

None

B

DeepSeek-V4-Flash

DeepSeek Official

chat_agent + create_branch

C

DeepSeek-V4-Flash

DeepSeek Official

None

Key Results

Metric

A (GLM-5.2)

B (DS + Tools)

C (DS Pure)

Total Output

38,888 chars

169,484 chars 🏆

57,565 chars

Total Time

705s

1,516s

239s 🏆

Total Tokens

28,459

210,482

29,184

Total Cost

¥0.75

¥0.21

¥0.06 🏆

Avg Output/Question

7,778 chars

33,897 chars (4.4x) 🏆

11,513 chars

Tool Calls

0

30 🏆

0

Cache Hit Rate

0%

up to 68% 🏆

0%

What We Found

  • With MCP tools, DeepSeek-V4-Flash produced 4.4x more detailed engineering solutions — including complete Go-style Raft consensus implementations, assembly-level RTOS scheduler code, and production-ready service mesh configurations

  • Tree-structured thinking (create_branch) enabled the model to explore 5-6 levels deep on complex problems, creating subtrees for architecture, implementation, testing, and verification

  • Cost comparison: GLM-5.2 costs 13x more than DeepSeek-V4-Flash pure thinking mode (¥0.75 vs ¥0.06) for comparable output quality. With tools enabled, DeepSeek-V4-Flash cost increased to ¥0.21 due to deeper exploration (3.7x more tokens), but remained 3.6x cheaper than GLM-5.2.

  • Cache hit rates reached 68% during multi-round tool calls, dramatically reducing effective input costs via DeepSeek's prefix caching

Full experiment results and data: results/comparison/report.md


License

MIT

Available Tools

1 tool
chat_agentA

调用独立、非思考模式的模型生成文本,用于延伸当前思维链。你可以传入任何需要独立完成且不受对话历史影响的子任务,包括但不限于:逻辑推理、发散联想、创意生成、优缺点分析、情景假设、步骤拆解、知识类比与迁移等。请确保 input_text 是一个完整、自包含的任务描述,明确任务类型与目标。通过调整 temperature(02)和 top_p(01)控制输出的确定性与多样性。对于需要精确复现的场景(如校验任务),建议设置 temperature 接近 0 并指定 seed。

ParametersJSON Schema
NameRequiredDescriptionDefault
input_textYes完整、自包含的任务描述。需明确任务类型(推理/联想/类比/创意/评估等)并提供所有必要的上下文信息,使工具无需依赖外部信息即可独立完成任务。示例:'推理:已知X=5, Y=12,请推导Z=X²+Y²的值,并给出计算步骤。'
system_promptNo可选的系统提示词,用于设定角色或行为约束。如:'你是一个严谨的数学校验员,只输出最终结果。'
temperatureNo采样温度 0.0-2.0。低值(0.0-0.3)=确定/精确,高值(0.7-2.0)=创造/发散
top_pNo核采样阈值 0.0-1.0。与temperature配合使用控制输出多样性
max_tokensNo最大输出token数。通过API参数在服务端控制,非本地硬截断
stopNo停止序列,遇到这些字符串时停止生成。默认 ['\n\n'] 防止非思考模型自动续写,思考模型可按需覆盖
seedNo随机种子(需 API 支持)。设置后配合低 temperature 可实现输出复现

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains that the model is non-thinking, that input must be self-contained, and how temperature/top_p/seed control output. However, it lacks disclosure of potential side effects like token cost or API latency, and without annotations the burden is higher. The description is adequate but could be more explicit about behavioral traits such as statelessness or repeatability guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that front-loads the purpose, then flows into usage guidelines and parameter details. Every sentence contributes meaning, but it is somewhat lengthy. For a tool with 7 parameters, this length is justified, and the structure is logical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the tool's full input contract and behavior adequately. It explains all parameters and their interplay. Missing are return value expectations (e.g., response format) and error handling, but for a straightforward generation tool, the provided context is sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 100%, the description adds significant value: it explains what constitutes a valid input_text with examples, provides intuition for temperature and top_p ranges, clarifies the stop default's purpose (preventing automatic continuation), and advises on seed use for reproducibility. This goes well beyond the schema field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool calls a non-thinking model to generate text for extending a thought chain, listing specific use cases like reasoning, creativity, and analysis. It uses a specific verb ('调用') and resource ('模型生成文本'), making the purpose unmistakable even without sibling tools for comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit context: it is for independent subtasks not influenced by conversation history. It lists many appropriate scenarios and implies that the tool should not be used for tasks requiring conversational memory, though it does not formally state exclusions. Given no sibling tools, this is clear and sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev1.0.0
    • First observedchat_agent

TDQS

A4.1/5.0
Disambiguation5/5

Only one tool exists, so there is no possibility of confusion between tools. The tool's purpose is clearly described.

Naming Consistency5/5

With a single tool, naming consistency is trivially maintained. The snake_case format is acceptable.

Tool Count3/5

The server has only one tool, which is at the lower end of reasonable scope. While a single general-purpose tool can be useful, it feels thin for a server named 'Thinking Agent MCP'.

Completeness2/5

The tool surface is severely incomplete for a thinking agent server. It only provides text generation without any supporting tools for memory, planning, or verification.

Maintenance

ActivityStale
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ScarletLilith/DeepSeekV4Flash_Thinking_TreeMCP'

If you have feedback or need assistance with the MCP directory API, please join our Discord server