Skip to main content
Glama

MCProbe

A stdio MCP server that audits other MCP servers over the live protocol. It connects to any MCP target (stdio or HTTP), lints every tool's schema for agent-usability, then actually calls the tools with deliberately broken inputs to see how the server handles them, and returns a 0–100 conformance score with a per-dimension breakdown rendered as Markdown.

The behavioral pass is the part that matters. Static schema audits tell you that a tool exists and looks reasonable. MCProbe then picks up a phone and dials each tool with missing_required, wrong_type, out_of_enum, and extra_garbage inputs — the same mistakes a language model will make on a bad day — and classifies the response. A server that returns a clean isError: true rejected the input correctly. A server that says "OK" to garbage (silently accepted it) or crashes the JSON-RPC transport both failed to reject it — and the Error Handling score is the fraction of bad inputs the server rejected cleanly.

Problem statement

The Model Context Protocol is new. Servers proliferate. Most ship with tool schemas that an agent can call, but few ship with tool schemas that an agent can call correctly: parameters are untyped, descriptions are missing, names are not snake_case, and a quick look at the code reveals that the handler is doing Number(x) / Number(y) with no guard at all.

The convention in the wider ecosystem is to ship a static schema audit that flags the obvious smells and then declare the server ready. The smells are real, but a static audit cannot tell you whether the server behaves: it cannot tell you that divide("x", "y") silently returns NaN, or that an extra unknown key is just stripped and ignored.

MCProbe does both, on a single connection:

  1. Static lint. Twelve rules over every tool's schema: missing or thin descriptions, duplicate or unusual names, an empty or non-object schema, untyped or undocumented parameters, and a server-wide rule for "I said I had tools but I have none."

  2. Behavioral fuzz. For each tool, the generator produces one valid case and at least three malformed variants, calls the target over the live JSON-RPC transport, and classifies the outcome as ok (the tool shrugged), toolError (graceful rejection), or protocolCrash (worst case). A malformed case that comes back without isError: true is flagged as silentlyAccepted — exactly the failure mode the linter cannot see.

  3. Scoring. The findings and the fuzz results are combined into a 0–100 score on four dimensions, mapped to an A–F grade, and rendered as a Markdown report the host (or a human) can read.

Related MCP server: owasp-mcp-scan

Hosted version — mcprobe.org

Don't want to install anything? mcprobe.org is the hosted version of this engine — paste an MCP server's URL in your browser and get the same graded report, no Node or setup required.

  • Free — 2 audits/day with a soft report (score, grade, dimension scores, finding counts).

  • Pro ($9.90 once, lifetime) — the full report (per-dimension reasons, every finding, the fuzz table, recommended fixes), 30 audits/day, saved history, the public gallery, Markdown export, and shareable links.

  • Local (stdio) servers — a Pro feature: run mcprobe push --stdio "…" --token <key> to audit a server on your machine and send the report to your account (see the CLI).

This engine stays MIT-licensed and free — the hosted app only adds accounts, persistence, the gallery, and those conveniences. Run it yourself for nothing, or pay once for the hosted experience.

Install

npm install
npm run build     # tsc -p tsconfig.json && tsc -p examples/demo-target/tsconfig.json

The build emits:

  • dist/index.js — the probe (run this as a stdio MCP server).

  • examples/demo-target/dist/index.js — a deliberately flawed MCP server used by the tests and the demo.

To launch the probe as a stdio MCP server so any host can talk to it:

npm start

No port, no daemon, no config file. The probe speaks JSON-RPC on stdin/stdout and writes operator logs to stderr.

Quickstart — audit any MCP server

Two ways to point MCProbe at a target. You only ever register MCProbe; it dials the target itself, so the target needs no setup.

Option 1 — from an MCP client (Claude Desktop, Cursor, any host)

Add MCProbe to your client's MCP config (use the absolute path to the built dist/index.js):

{
  "mcpServers": {
    "mcprobe": {
      "command": "node",
      "args": ["/absolute/path/to/mcprobe/dist/index.js"]
    }
  }
}

Then ask in plain English:

Use mcprobe to audit https://docs.base.org/mcp over http — connect, then run a full report with fuzz and show me the score.

The host calls probe_connect then probe_report for you. MCProbe also advertises server instructions, so the model is told the flow on connect — no need to memorise the tool names.

Option 2 — the mcprobe CLI (no host)

Get the project and build it, then audit any server straight from the terminal:

git clone https://github.com/alitiknazoglu/mcprobe
cd mcprobe && npm install && npm run build

# audit an HTTP server
node dist/index.js audit https://docs.base.org/mcp --fuzz

# audit a LOCAL stdio server (no URL — the `npx some-server` style)
node dist/index.js audit --stdio "npx @acme/my-mcp-server" --fuzz

# audit a server that requires credentials
node dist/index.js audit https://api.acme.com/mcp --bearer sk_live_xxx --fuzz
node dist/index.js audit https://api.acme.com/mcp --header "X-API-Key: xxx"

--bearer/--header authenticate you to the server being audited, and are sent on every request to it — so only point them at a server you trust. Set MCPROBE_TARGET_TOKEN instead of --bearer to keep the key out of your shell history. (Don't confuse these with push --token, which authenticates you to MCProbe; passing that as --bearer would hand your MCProbe key to a third party.) A stdio server takes its credentials through its own environment instead.

(After npm install -g . or npm link, the command is just mcprobe audit ….)

It prints the full Markdown report to stdout — add --json for a machine-readable report (what the GitHub Action and other tooling consume). --fuzz also calls each tool with malformed input to score Error Handling & Liveness; tools the target marks destructiveHint: true are skipped unless you add --fuzz-destructive, so a default run is safe even against servers you don't control. Omit --fuzz for a read-only static audit (metadata + schema quality only).

Save an audit to your account. push runs the same audit and uploads the report to an ingest endpoint (default https://mcprobe.org/api/ingest) with a bearer token:

node dist/index.js push --stdio "npx @acme/my-mcp-server" --fuzz --token mcp_xxx

The token comes from your mcprobe.org profile; --to <url> (or MCPROBE_API) points it at a different endpoint. Run mcprobe help for all flags.

Audit in CI (GitHub Action)

Gate your MCP server on every push — audit it and fail the build if its conformance grade drops. Free and self-contained (it runs the open-source engine on your own runner; no account required):

# .github/workflows/mcprobe.yml
name: MCP audit
on: [push, pull_request]
jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: alitiknazoglu/mcprobe@v1
        with:
          url: https://your-server.example.com/mcp
          fuzz: true          # behavioral testing (call tools with bad input)
          min-score: "75"     # fail the job below this (A≥90 B≥75 C≥60 D≥40); omit to report only

The step prints the score to the job summary and exposes score / grade outputs. It also writes a full mcprobe-report.json you can upload as an artifact. Leave off min-score to report without ever failing the build.

Upload runs to your dashboard (Pro): add one line — a token: (your mcprobe.org Pro key, stored as a GitHub secret). The audit stays the same; the run is also uploaded to your history/dashboard on mcprobe.org.

      - uses: alitiknazoglu/mcprobe@v1
        with:
          url: https://your-server.example.com/mcp
          min-score: "75"
          token: ${{ secrets.MCPROBE_TOKEN }}   # ← only new line; uploads to your dashboard

The audit itself is always free and local; the hosted tracking (history, gallery, badge) is the Pro tier — see Hosted version.

Agent skill

This repo ships an agent skill so your coding agent knows how to drive MCProbe on its own — just say "audit this MCP server" and it runs the right probe_* tools or mcprobe CLI command and explains the score. To install it, copy the folder into your agent's skills directory:

# Claude Code (project- or user-level)
cp -r .agents/skills/mcp-audit /path/to/your/project/.claude/skills/

# Other agents that use the open skills format (Codex, Opencode, Cursor, …)
cp -r .agents/skills/mcp-audit /path/to/your/project/.agents/skills/

It's a single SKILL.md — the same file works in either location.

The six probe_* tools

MCProbe registers four core tools and two optional helpers. The core four cover the full lint → fuzz → score pipeline; the two helpers cover the everyday ergonomics of managing connections.

Tool

Purpose

Returns

probe_connect

Open a connection to a target.

{ connectionId, name, version, capabilities, counts, defaultConnectionId }

probe_lint

Run the 12 lint rules over the target's cached tool summaries.

{ connectionId, server, findings, summary }

probe_fuzz

Generate valid + malformed inputs per tool, call each, classify the outcome. Skips destructive tools by default.

{ connectionId, server, results, coverage, summary }

probe_report

Run lint (and fuzz when requested), score, render Markdown.

{ connectionId, server, overall, grade, dimensions, coverage, findings, fuzz, markdown }

probe_list

(optional) Enumerate the target's tools.

{ connectionId, server, tools }

probe_disconnect

(optional) Close one connection (by id) or every connection.

{ removed, remaining, defaultConnectionId }

All tools default to the most recently opened connection when connectionId is omitted, so a single-target audit is a three-call sequence: probe_connectprobe_reportprobe_disconnect.

Every tool also declares MCP annotations so a host can reason about side effects before calling: probe_lint and probe_list are readOnlyHint: true, while probe_fuzz is destructiveHint: true (it invokes the target's tools), and the tools that reach a target (probe_connect, probe_fuzz, probe_report) set openWorldHint: true. MCProbe audits other servers for agent-usability, so it declares these hints on its own tools too.

probe_connect

Two transports: stdio (spawns a child process) and http (speaks the streamable HTTP transport, with SSE fallback). For stdio, command is required; for http, url is required. The target's initialize handshake is run synchronously, the server's identity and capabilities are cached, and a stable connectionId is returned.

probe_lint

A pure pass over the connection's cached tool summaries — no extra round-trip. Each finding carries a stable code, a severity (error, warning, info), a human-readable message, a location ({ tool, param? }), and a hint with a concrete fix.

The twelve rules are:

Code

Severity

What it catches

tool.missing_description

error

A tool with no description at all.

tool.thin_description

warning

A description under 12 characters.

tool.duplicate_name

error

Two tools registered with the same name.

tool.unusual_name

warning

A name that is not snake_case or kebab-case.

tool.no_input_schema

warning

An empty or missing inputSchema.

tool.no_annotations

info

A tool that declares no MCP annotations (readOnlyHint, destructiveHint, etc.).

schema.invalid

error

A schema that fails to compile (Ajv).

schema.root_not_object

warning

A root type that is not object.

schema.no_required

info

Properties declared but no required array.

param.untyped

warning

A property with no type/enum/const/oneOf.

param.missing_description

warning

A property with no description.

server.no_tools

warning

The server claims tools but registers none.

probe_fuzz

For every tool (capped at maxTools, default 10), the generator emits one valid case and at least three malformed variants:

  • missing_required:<field> — drop each required field in turn.

  • wrong_type:<field> — replace each typed field with a value of a different primitive type.

  • out_of_enum:<field> — for enum or const fields, send a value the schema forbids.

  • extra_garbage — append a sentinel key to the valid args.

Each case is sent to the target over the live JSON-RPC transport. The classifier assigns one of three outcomes:

Outcome

Meaning

ok

The target returned a result with isError: false. For a malformed case this is silentlyAccepted: true; for a valid case with no usable content it is emptySuccess: true, and a valid success that breaks the tool's declared outputSchema is outputSchemaViolation: true.

toolError

The target returned a result with isError: true (graceful rejection).

protocolCrash

The call rejected or the transport closed.

Hallucinated success (emptySuccess). A valid call that returns success but with an empty / contentless result — e.g. a write tool that answers 200 with an empty body and never persists anything. The agent reads "done" while nothing happened. MCProbe flags this on the critical line, drops Liveness credit for that call (it isn't a real success), and recommends returning a confirmation payload. This is the "the agent said done, nothing happened" bug. (A tool that returns valid structuredContent but empty text content is not flagged — the structured payload is a real result.)

Output-schema conformance (outputSchemaViolation). When a tool declares an outputSchema (MCP structured output), MCProbe checks that a valid call actually honors it — the success must return structuredContent that validates against the declared schema. A tool that advertises an output contract and then returns no structured content, or a payload that doesn't match, is flagged (and reported as an output-schema violation rather than a misleading "protocol crash"). Not credited in Liveness — a success that breaks its own contract isn't a real success.

Dry-run safety. By default, tools annotated destructiveHint: true are not fuzzed — so pointing MCProbe at a server you don't control can't trigger a real destructive action (e.g. a delete_file tool). Pass fuzzDestructive: true to override. probe_fuzz (and the report) return a coverage summary listing how many tools were fuzzed and which were skipped (as destructive, or over the maxTools cap).

probe_report

The convenience entry point. Calls probe_lint (always) and probe_fuzz (when fuzz: true), scores the result on the four dimensions described below, and returns the structured ConformanceReport and a rendered Markdown string. The Markdown is the canonical payload; downstream tools that need the numbers can pull them out of the structured fields.

Scoring model — four dimensions

The overall 0–100 score is the mean of the measured dimensions. Dimensions that were not measured (e.g. the two behavioral ones when fuzz: false, or when every tool was skipped) are reported as "not measured" and excluded from the average rather than penalized with a fake value. This is what lets a static audit of a clean server still score 100/100.

The two static dimensions are subtractive (start at 10, lose points per finding). The two behavioral dimensions are normalized rates, so a score is comparable across servers of different sizes — and the fuzz cases are partitioned by kind (malformed → Error Handling, valid → Liveness) so no outcome is ever counted twice.

Letter grades: A ≥ 90, B ≥ 75, C ≥ 60, D ≥ 40, F < 40.

Dimension

Always measured?

What it captures

Metadata & Documentation

yes

Server identity (name, version), advertised capabilities, presence of instructions (+1 bonus).

Schema Quality

yes

Rate over tools: findings are weighted (1 per error, 0.5 per warning, 0.25 per info), summed per tool and capped at 2 each, then averaged across every tool the server advertises. A server isn't penalized for having more tools — 50 tools with one minor nit each score the same as 5.

Error Handling

only with fuzz: true

Rate over malformed cases: 10 × (gracefully-rejected / total malformed). A silent accept (garbage let through) or a protocol crash both count as failed rejections.

Liveness & Performance

only with fuzz: true

Rate over valid cases: 10 × (successful / total valid), minus 0.5 per 100ms that the valid-call p50 latency exceeds a 200ms target.

The per-dimension reasons and counts are emitted in the Markdown report so the score is auditable by a human. When fuzzing runs, the report header also shows two extra lines:

  • a Coverage line (how many tools were fuzzed, and which were skipped as destructive or over the maxTools cap); and

  • a critical-issues callout — a flag, not a second score — hoisting the dangerous findings to the top, e.g. ⚠ Critical: 4 tool(s) silently accept malformed input (…); 1 protocol crash(es), or ✓ No critical behavioral issues when there are none. The normalized scores are unchanged; this just makes the scary stuff visible above the fold.

The report ends with a Recommended fixes section: a prioritized to-do list (worst severity first) that turns each finding into a concrete action — the fix hint plus the exact tools/parameters it affects — followed by behavioral fixes for tools that silently accept input or crash. So the report is a prescription, not just a diagnosis. A clean server gets "Nothing to fix — this server passes every check."

30-second demo

The probe ships with a deliberately flawed demo target at examples/demo-target/ and a smoke script that runs the full probe_report pipeline against it. From a clean clone:

npm install
npm run build
node scripts/smoke-report.mjs

The script spawns the probe as a stdio MCP server, opens a connection to the demo target, calls probe_report with fuzz: true, and prints the Markdown report to stdout. The demo target is wired to fail loudly: greet has no description, divide returns NaN on bad input, set_mode has a thin description, and well_behaved is the only tool with a clean, validated schema. The report will show a low overall score with concrete findings, a coverage line, a critical-issues callout, and a fuzz table that classifies the broken cases.

For an interactive tour, the official MCP inspector works as a host against the built probe:

npx @modelcontextprotocol/inspector node dist/index.js

The inspector UI lists the six probe_* tools; calling them manually is a good way to see the request/response shape.

External server example

For a full probe_connectprobe_reportprobe_disconnect walkthrough as an AI agent would run it (natural-language request, the JSON tool calls, and the rendered report), see examples/agent-usage.md.

The probe is not coupled to the demo target. To audit any other MCP server, swap the command/args in probe_connect:

// tool call: probe_connect
{
  "transport": "stdio",
  "command": "npx",
  "args": ["-y", "@modelcontextprotocol/server-filesystem@latest", "/tmp"]
}

The probe runs the initialize handshake against the spawned process, caches its tools, and is ready for probe_lint / probe_fuzz / probe_report. The same pattern works for HTTP targets: pass transport: "http" and a url instead.

A real transcript of this audit (run against @modelcontextprotocol/server-filesystem@latest and saved to examples/transcripts/external-server.md) is included in the repository. The script that produced it is scripts/external-audit.mjs. A self-audit (a second copy of the probe scoring the first) lives at examples/transcripts/self-audit.md.

Use as a library

Besides the MCP server, MCProbe exposes its audit pipeline as functions for embedding in your own backend:

import { auditUrl, auditStdio, softenReport, renderReport } from "mcprobe/audit";

// HTTP server (URL in, report out — never spawns a process):
const report = await auditUrl("https://example.com/mcp", { fuzz: false });

// A server that requires credentials — headers ride along on every request
// (connect *and* tool calls), so the whole audit works, fuzzing included:
const gated = await auditUrl("https://api.acme.com/mcp", {
  fuzz: true,
  headers: { Authorization: "Bearer sk_live_xxx" },
});

// Local stdio server (spawns the subprocess — only run commands you trust):
const local = await auditStdio("npx", { args: ["@acme/my-mcp-server"], fuzz: true });

console.log(report.overall, report.grade);  // structured ConformanceReport
console.log(renderReport(report));          // or the Markdown
const teaser = softenReport(report);        // a trimmed view (scores, no detail)

Both default to a static, read-only audit (fuzz: false); pass fuzz: true to also run the behavioral fuzzer (destructive tools are skipped unless fuzzDestructive: true). auditUrl is HTTP-only and side-effect-free, ideal for a hosted backend; auditStdio launches a local subprocess, so use it only for servers you trust (CLIs, your own machine). softenReport is handy for a free/preview tier — it keeps the scores and counts but withholds the reasons, full findings, fuzz table, and recommended fixes.

This is exactly how the hosted app at mcprobe.org is built on top of the engine.

Architecture

MCProbe plays two roles at once: it is a stdio MCP server to its host, and an MCP client to whatever it is auditing. The split mirrors the source layout.

+-------------------------------------------------+
| any MCP client over stdio:                      |
| Claude Code, an IDE, an agent, or a node script |
+-------------------------------------------------+
                          |
                          |  stdio JSON-RPC  (stdin / stdout)
                          v
+--------------------------------------------------+
|  MCProbe  -  one stdio MCP server                |
|                                                  |
|  src/index.ts       registers the probe_* tools  |
|      |  then calls the pure modules:             |
|      +--> src/schema-lint   (12 lint rules)      |
|      +--> src/fuzz          (case generator)     |
|      +--> src/conformance   (4-dimension score)  |
|      +--> src/report        (markdown renderer)  |
|      |                                           |
|      v                                           |
|  src/target-client  (outbound MCP client)        |
+--------------------------------------------------+
                          |
                          |  stdio / http JSON-RPC
                          v
              +---------------------+
              |  target MCP server  |
              +---------------------+

The top box is whatever drives MCProbe over stdio — a full host like Claude Code, or a plain node script (the scripts/*.mjs drivers and the Quickstart's audit.mjs are exactly this; no host required). It talks only to MCProbe; MCProbe's src/target-client then dials the audited server over stdio or http. The probe sits in the middle — a server to its caller, a client to its target.

Module

Role

I/O?

src/types.ts

Shared Finding, FuzzResult, DimensionScore, ConformanceReport types.

none

src/target-client.ts

Outbound MCP client, ConnectionRegistry, callTool wrapper that catches transport errors.

yes — spawns / dials

src/schema-lint.ts

The 12 lint rules. Pure: no I/O, deterministic ordering.

none

src/fuzz.ts

Case generator + runner + summarizeFuzz histogram. Generator is pure; runner threads through a caller-supplied call fn so it stays unit-testable.

none on the generator; the runner calls the target

src/conformance.ts

Per-dimension scoring + rollup. Pure.

none

src/report.ts

Pure Markdown renderer. Same input → same output every run.

none

src/index.ts

McpServer, registers the six probe_* tools, routes them to the pure modules.

yes — owns the stdio transport

The four pure modules (schema-lint, fuzz generator, conformance, report) are deliberately side-effect-free so the vitest suite can exercise them in milliseconds without spawning a target. The integration test in tests/demo-target.test.ts is the only piece that touches a live process; it is the smallest test that proves the build artifact loads over the real protocol.

Limitations

  • The four runtime dependencies are frozen. @modelcontextprotocol/sdk, ajv, ajv-formats, zod. The probe deliberately does not depend on any CLI framework, HTTP server, or transport library beyond what the SDK already exposes. Adding a runtime dependency is an explicit change to the spec.

  • The probe is a stdio MCP server, full stop. It does not expose an HTTP endpoint. Run it as a subprocess of your host.

  • Only static credentials are supported. A server behind an API key or bearer token can be audited with --bearer / --header (or the headers option in the library). A server that requires a full OAuth login flow — an authorization-code/PKCE dance with dynamic client registration — cannot be audited yet, because MCProbe has no way to complete that flow on your behalf.

  • The fuzzer is shallow, not adversarial. It exercises the surface documented by the tool's inputSchema; it does not attempt to discover server-side bugs that are out of band of the tool contract. The point of MCProbe is conformance, not general-purpose server fuzzing.

  • The scoring is dimension-local. A perfect score on one dimension does not rescue a failure on another. The static dimensions are subtractive; the behavioral dimensions are normalized rates. The four dimensions are weighted equally when measured.

  • Dry-run skips destructive tools. By default a fuzz run does not exercise tools annotated destructiveHint: true; they show up in the coverage summary as skipped. A target that doesn't annotate a destructive tool will still be fuzzed — annotations are the only signal MCProbe has. Pass fuzzDestructive: true to fuzz everything.

  • Behavioral scores need a real protocol round-trip. When fuzz: false is passed to probe_report, the Error Handling and Liveness & Performance dimensions are reported as "not measured" and excluded from the rollup. A "lint-only" audit can still score 100/100 on a clean server, but it cannot tell you whether the server would survive a bad input.

  • Tooling is four cores + two helpers, no more. The spec pins the surface area. Adding a probe_* tool is an explicit change to the spec.

  • The optional helpers are still required at startup. The McpServer is constructed with the tools capability only; it does not advertise resources or prompts. The probe itself is an audit tool, not a content server.

Available Tools

6 tools
probe_connectA

Open a connection to an MCP target server (stdio or HTTP) and return a connectionId plus the server's name, version, capabilities, and counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
envNostdio: extra environment variables layered on top of the parent process env.
urlNohttp: URL of the target MCP server. Required when transport='http'.
argsNostdio: arguments passed to the spawned process.
commandNostdio: executable to spawn (e.g. 'node'). Required when transport='stdio'.
transportYesTransport to use. 'stdio' spawns a child process; 'http' speaks the streamable HTTP transport.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this tool is not read-only (readOnlyHint=false), and the description states it 'opens a connection' implying persistence. However, it does not elaborate on side effects, resource management, or behavioral differences between transport types beyond parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the key action and output, with no unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, no output schema), the description covers the main output but lacks details on error handling, timeouts, or cleanup. Still, it provides sufficient context for basic use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameter details are fully documented in the schema. The description adds no further semantic information beyond referencing the transport types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: opening a connection to an MCP target server and returning specific information (connectionId, server name, version, etc.). It distinguishes well from siblings like probe_disconnect, probe_lint, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing to connect to an MCP server, but does not explicitly state when to use or avoid this tool compared to alternatives. No prerequisites or limitations are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_disconnectA
Idempotent

Close a single connection (by id) or all connections if id is omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoConnection id returned by probe_connect. Omit to close every connection.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate write operation (readOnlyHint=false) and idempotency. The description adds critical detail: behavior changes based on id presence (single vs all). No contradictions, but could mention potential errors (e.g., invalid ID) or safety concerns with closing all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that front-loads the action and covers both cases. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description provides essential behavior. It could be enhanced with error handling or return value hints, but is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already covers the single parameter with full description (connection ID and omission behavior). The tool description adds no extra meaning beyond the schema. High schema coverage justifies baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool closes connections, specifying the two modes: single connection by ID or all if omitted. It effectively distinguishes from siblings like probe_connect (opening) and probe_list (listing).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (closing connections) and shows two scenarios. It doesn't explicitly state when not to use it or list alternatives, but the sibling names provide differentiation. Adequate for a simple tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_fuzzA
Destructive

Generate one valid and several malformed inputs per target tool, call each, and record the outcome (ok, toolError, protocolCrash), whether malformed inputs were silently accepted, and call latency. Tools annotated destructiveHint:true are skipped by default (set fuzzDestructive to include them). Returns a coverage summary of which tools were fuzzed vs skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxToolsNoCap on the number of tools to fuzz. Defaults to 10.
connectionIdNoIdentifier returned by probe_connect. Defaults to the most recent connection.
fuzzDestructiveNoAlso fuzz tools annotated destructiveHint:true. Default false (the dry-run safety guard) so fuzzing an untrusted target can't trigger a destructive action.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description discloses detailed behavior: generating one valid and several malformed inputs, recording outcomes, latency, and silent acceptance. It explicitly notes destructive tools skipped by default and the safety guard, adding significant context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense in a single sentence, which is efficient but slightly run-on. It front-loads key operations and outcomes, earning its content without significant waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (fuzzing, safety, multiple outputs), the description covers input generation, calling, outcome recording, default skip behavior, and return value (coverage summary). Minor ambiguity about whether detailed outcomes are returned, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the description adds value by explaining the safety role of fuzzDestructive and functional role of maxTools and connectionId. It provides context not present in the schema, justifying a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: generate valid and malformed inputs per target tool, call each, and record outcomes. It specifies verb+resource with precise scope, and distinguishes well from sibling tools like probe_connect and probe_lint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains default behavior (skipping destructive tools) and provides guidance on when to use fuzzDestructive. It implies safety context but does not explicitly state when not to use or describe alternatives beyond the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_lintA
Read-onlyIdempotent

Run the lint rules over the target's tool schemas and return a list of findings with stable codes, severities, locations, and fix hints.

ParametersJSON Schema
NameRequiredDescriptionDefault
connectionIdNoIdentifier returned by probe_connect. Defaults to the most recent connection if omitted.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint, idempotentHint, openWorldHint. Description adds that findings have stable codes, severities, locations, and fix hints, clarifying the output characteristics beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with action, no fluff. Efficiently communicates purpose and output for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Describes return value (list of findings with codes, severities, locations, fix hints) despite no output schema. Covers essential behavioral aspects for a tool with one optional parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already describes the parameter (connectionId) fully. Description adds no additional semantic information beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states action 'Run lint rules over tool schemas' and output 'findings with stable codes, severities, locations, and fix hints'. Differentiates from siblings (connect, fuzz, report, list, disconnect) through naming and description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage after connecting (via connectionId) but does not explicitly state when to use vs alternatives, prerequisites, or when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_listA
Read-onlyIdempotent

Enumerate the target's tools (name, description, input schema) using the default or a specific connection.

ParametersJSON Schema
NameRequiredDescriptionDefault
connectionIdNoIdentifier returned by probe_connect. Defaults to the most recent connection.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint. The description adds value by detailing the output (name, description, input schema) and the connection parameter behavior (default vs specific). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence of 15 words, front-loaded with the action and resource. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While there is no output schema, the description adequately conveys the return information (name, description, input schema). It covers the main usage scenario but omits edge cases like missing connections.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description's mention of 'default or a specific connection' reinforces the schema's parameter description but adds no new meaning beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the verb 'Enumerate' with resource 'tools' and specifies the returned structure (name, description, input schema). It clearly distinguishes from sibling tools like probe_connect and probe_lint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly indicates use after a connection, but provides no explicit when-to-use or when-not-to-use guidance. No alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_reportA

Run introspect + lint (and fuzz when requested) against the target, score the result on four dimensions, and return a Markdown report with the overall score, letter grade, per-dimension breakdown, findings, and fuzz table.

ParametersJSON Schema
NameRequiredDescriptionDefault
fuzzNoWhen true, run the behavioral fuzzer before scoring. Default false; only static dimensions are measured when omitted.
maxToolsNoForwarded to probe_fuzz when fuzz=true. Defaults to 10.
connectionIdNoIdentifier returned by probe_connect. Defaults to the most recent connection.
fuzzDestructiveNoForwarded to probe_fuzz when fuzz=true. Also fuzz tools annotated destructiveHint:true (default false — the dry-run safety guard).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false and openWorldHint=true, indicating possible side effects. The description adds context about running optional fuzz and defaults, but does not elaborate on behavioral traits like destruction, auth needs, or rate limits beyond what annotations and schema already convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence efficiently lists all actions and output components with no filler. Every part earns its place; it is front-loaded with the main verb and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description fully explains the return value (Markdown report with overall score, letter grade, per-dimension breakdown, findings, fuzz table). It also covers optional fuzz behavior and defaults, making the tool's behavior complete for selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds minor context (e.g., fuzz triggers behavioral fuzzer) but does not significantly enhance parameter meaning beyond the schema's own descriptions and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('run introspect + lint + fuzz') and the resource ('the target'), specifies the output ('Markdown report with score, grade, breakdown, findings, fuzz table'), and distinguishes from siblings by combining multiple operations into one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that fuzz is optional via 'when requested' and gives a default for fuzz. However, it does not explicitly say when to use this tool versus individual sibling tools (e.g., probe_lint, probe_fuzz) or provide any exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 6 tool updatesv0.1.0
    • First observedprobe_connect
    • First observedprobe_disconnect
    • First observedprobe_fuzz
    • First observedprobe_lint
    • First observedprobe_list
    • First observedprobe_report

TDQS

A4.3/5.0
Disambiguation5/5

Each tool has a distinct purpose: connect, lint, fuzz, report, list, disconnect. No overlap or ambiguity in their functions.

Naming Consistency5/5

All tools follow the consistent verb_noun pattern with 'probe_' prefix (e.g., probe_connect, probe_lint), ensuring predictability.

Tool Count5/5

Six tools cover the necessary operations for probing MCP servers (connect, list, lint, fuzz, report, disconnect) without being excessive or insufficient.

Completeness5/5

The tool set covers the full lifecycle: connection, introspection (list, report), linting, fuzzing, and disconnection. No obvious gaps for the stated domain.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Passive security scanner that audits a running MCP server against the OWASP MCP Top 10 and grades it A-F. Read-only static analysis of the advertised tools, prompts and resources with console/JSON/SARIF output, and it also runs as an MCP server itself.
    85
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Audits MCP server configurations for security risks including capability inventory, SSRF, prompt injection, and drift detection. Works in read-only mode and can also be used as an MCP server to let AI agents audit their own attack surface.
    4
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/alitiknazoglu/mcprobe'

If you have feedback or need assistance with the MCP directory API, please join our Discord server