Skip to main content
Glama
deckflow

DeckProbe MCP Server

Official
by deckflow

DeckProbe MCP Server

Let an agent ask what's inside a PDF, Office, or iWork file — without opening it.

CI npm License: MIT

Install · Tools · Configuration · Security · How it works · DeckProbe

An MCP server that exposes DeckProbeffprobe for documents — as four typed tools. Ask for page counts, slide counts, metadata, encryption and macro signals, structure, or integrity, and get back bounded, deterministic JSON with confidence, evidence, and measured I/O cost.

Nothing is rendered, no macro runs, no external reference is followed, and no network connection is opened. It is safe to point at untrusted files.

// probe { "path": "deck.pptx", "targets": ["slide_count"], "view": "values" }
{
  "schema_version": 2,
  "status": "ok",
  "driver": { "id": "powerpoint", "profile": "pptx" },
  "values": { "powerpoint.slide_count": 31 },
  "view": "values"
}

Install

Nothing to install ahead of time — npx fetches the server and the engine together.

Claude Code

claude mcp add deckprobe -- npx -y @deckflow/deckprobe-mcp

Claude Desktop, Cursor, VS Code, Zed, and anything else reading mcpServers

{
  "mcpServers": {
    "deckprobe": {
      "command": "npx",
      "args": ["-y", "@deckflow/deckprobe-mcp"]
    }
  }
}

For a pinned install, npm install -g @deckflow/deckprobe-mcp and use deckprobe-mcp as the command.

Requires Node.js 20 or newer. The engine binary arrives as a per-platform optional dependency for macOS, Linux (glibc and musl), and Windows on x86-64 and ARM64; anywhere else the server falls back to the same engine compiled to WebAssembly, so npx works wherever Node does.

Related MCP server: flexorch-mcp

Tools

Tool

Use it for

probe

Everything about one document

probe_batch

Inventory or triage many documents in one call

list_formats

Which formats are supported, and where support stops

list_targets

The exact target names a format offers

There is also one resource, deckprobe://schema, carrying the report JSON Schema bundled with the running engine.

probe

{
  "path": "reports/q3.pptx",
  "targets": ["@summary", "@security"],  // presets, short names, or canonical names
  "level": "metadata",                   // header | metadata | deep
  "min_confidence": "high",              // low | medium | high | exact
  "target_confidence": { "slide_count": "exact" },
  "view": "report",                      // report | values
  "budget": { "max_physical_bytes": 8388608, "timeout_ms": 1000 }
}

targets accepts short names (slide_count), canonical names (powerpoint.slide_count), and presets:

Preset

Expands to

@header

Container identity only — format, size, extension match, encryption flag

@summary

Identity, common metadata, and primary structure

@security

Encryption, macros, signatures, external references, active content

@structure

Format-owned counts, names, and dimensions

@assets

Images, media, previews, fonts, embedded objects

@quality

Integrity, repair, extension match, conformance

@format

Every format-specific target at the active level

@all

Everything available at the active level

@summary deliberately omits statistics that need a full-file read. A PDF's page_count is the notable case — ask for it explicitly.

probe_batch

{ "paths": ["a.pdf", "b.pptx", "c.xlsx"], "targets": ["@security"] }

One engine process handles the whole batch. Results come back in input order, each with its own report or its own error, so one bad file never spoils the run. Defaults to the compact values view. Literal paths only — expand globs yourself.

list_formats and list_targets

list_targets takes a format (pdf, docx, xlsx, pptx, doc, xls, ppt, key, numbers, pages) and returns each target's aliases, description, value type, minimum level, cost class, and selector membership. Pass detail: "full" for the engine's complete report, including per-target JSON Schema fragments and expanded selector lists.

Both are cached for the lifetime of the server process.

Reading a report

The tool result is the engine's own schema-v2 envelope, unmodified. Two things are worth knowing before consuming it:

  • status: "partial" is not a failure. It means at least one requested target could not be resolved at the requested confidence. It is named in execution.unresolved_targets, and every other result still stands.

  • confidence_score is a fixed constant per label (0.4, 0.7, 0.95, 1.0), not a calibrated probability. 0.95 does not mean the value is right 95% of the time.

Only results with status resolved or estimated carry a value. unknown is common and usually means the document simply does not record that fact.

A failing call returns isError with the engine's error envelope plus one line saying what to do about it. Failures the server itself raises before the engine runs — a missing path, a directory, a path outside the allow-list, an exceeded deadline — use the same envelope shape with an MCP_-prefixed code and origin: "mcp-server".

Configuration

Every setting is an environment variable, set in your client's MCP config. All are optional.

Variable

Default

Meaning

DECKPROBE_MCP_BIN

Engine binary to use instead of the bundled one

DECKPROBE_MCP_ROOTS

unrestricted

Allowed directories, separated like PATH

DECKPROBE_MCP_TIMEOUT_MS

30000

Hard per-call deadline on an engine process

DECKPROBE_MCP_MAX_CONCURRENCY

4

Concurrent engine processes

DECKPROBE_MCP_MAX_BATCH

64

Paths accepted by one probe_batch call

{
  "deckprobe": {
    "command": "npx",
    "args": ["-y", "@deckflow/deckprobe-mcp"],
    "env": { "DECKPROBE_MCP_ROOTS": "/Users/me/Documents:/Users/me/Downloads" }
  }
}

Security

DeckProbe is built for untrusted input: bounded parsing, no renderer, no macro interpreter, no external-reference resolution, and no network access. This server adds two things on top.

  • Process isolation and a hard deadline. Each probe runs in its own short-lived process, killed if it outruns DECKPROBE_MCP_TIMEOUT_MS.

  • An optional read allow-list. DECKPROBE_MCP_ROOTS pins the reachable tree; paths are symlink-resolved before the check, so a link cannot step around it. The default is unrestricted, matching the CLI the user could run themselves — set it for shared or automated deployments.

Reports describe a document (metadata, counts, signals) rather than reproducing its contents. Note that report values such as a document title are still attacker-controlled strings: the server passes them through as JSON data and never interpolates them into instructions, and a consumer should treat them the same way.

Report a vulnerability privately as described in SECURITY.md.

How it works

MCP client
    │  JSON-RPC over stdio
    ▼
deckprobe-mcp ── validates arguments, resolves the path, maps the result
    │  argv + stdout (one process per probe, or one --jsonl process per batch)
    ▼
DeckProbe engine ── plans the cheapest paths that answer the request

The server spawns the native DeckProbe CLI rather than calling the WebAssembly build. The CLI reads only the byte ranges a probe plan needs, where the WebAssembly path holds the whole file in memory, and a separate OS process both isolates untrusted parsing and can be killed outright. The engine is chosen in this order:

  1. DECKPROBE_MCP_BIN

  2. the binary that ships with this package's @deckflow/deckprobe dependency

  3. deckprobe on PATH

  4. the bundled WebAssembly engine

The resolved engine is logged to stderr at startup. stdout belongs to the MCP transport and carries nothing else.

MCP server or agent skill?

DeckProbe also ships an Agent Skill that teaches a shell-capable agent to use the CLI directly. Both teach the same vocabulary. Use the skill when the agent has a shell and you want the CLI's full surface; use this server when it does not, or when you want typed arguments validated before the engine ever runs.

Development

npm install
npm test          # typecheck, lint, build, and the full suite
npm run test:watch

Contributions are welcome — see CONTRIBUTING.md. The design rationale, including the alternatives that were rejected, is in docs/rfc.md.

License

MIT. See LICENSE.

Available Tools

4 tools
list_formatsList supported formatsA
Read-onlyIdempotent

List the document formats DeckProbe can inspect: drivers, the extensions each one handles, and where support stops. Call this when you are unsure whether a file type is supported at all.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, non-destructive behavior. The description adds valuable context about the exact output content (drivers, extensions, support limitations), which goes beyond the annotations. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the primary purpose and scope front-loaded. It avoids redundancy and every sentence earns its place—the first states what it lists, the second when to call it. No unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless list tool with no output schema, the description fully explains what the tool returns (drivers, extensions, and support limits). There is nothing missing for an agent to correctly invoke it and interpret the result. Complexity is low, so this is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is fully covered (100% by default). There is nothing to add about parameters; the description doesn't need to explain any. The baseline for zero parameters is 4, and no additional info is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a clear resource ('document formats DeckProbe can inspect'), and specifies the content (drivers, extensions, and where support stops). It effectively distinguishes itself from sibling tools like probe and list_targets by focusing on format capabilities rather than probing or target listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'Call this when you are unsure whether a file type is supported at all.' It doesn't mention alternatives, but the trigger condition is clear and implies that if you have a specific file, you would use probe instead. It could be improved by naming the sibling tools explicitly, but the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_targetsList targets for a formatA
Read-onlyIdempotent

List every target a format supports, with its aliases, value type, minimum probe level, cost class, and which @selectors include it. Call this before naming a target you have not already seen — probe rejects an unknown one rather than guessing.

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNocompact (default): target names, aliases, descriptions, levels, and selector membership. full: the complete report, including each target's JSON Schema fragment and every selector's expanded member list.
formatYesA format profile from list_formats, such as "pdf", "docx", "xlsx", "pptx", "doc", "xls", "ppt", "key", "numbers", or "pages".

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, covering safety. The description adds behavioral context by spelling out returned fields and the probe rejection behavior, which matters for planning calls. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences front-load the core purpose and output contents, then give one practical usage rule. No filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a rich annotation set, fully documented parameters, and a description that names the output fields and the prerequisite call, an agent has enough to invoke the tool correctly even with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already explains format and detail fully, including enumerations and examples. The description adds no new parameter-level semantics beyond naming the report contents, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (list) and resource (targets for a given format), enumerates exactly what is included (aliases, value type, minimum probe level, cost class, selector membership), and contrasts with probe by stating probe rejects unknown targets. This distinguishes it from sibling tools list_formats and probe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs when to call: before naming a target not already seen, because probe rejects unknowns rather than guessing. It also ties the format parameter to list_formats, implying the prerequisite and differentiating from list_formats.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probeProbe a documentA
Read-onlyIdempotent

Read facts about one local PDF, Microsoft Office, or Apple iWork document without opening or rendering it: page, slide, and sheet counts; title, author, and dates; encryption, macro, signature, and active-content signals; structure; embedded assets; and integrity. Handles .pdf, .docx/.xlsx/.pptx, legacy .doc/.xls/.ppt, and modern .key/.numbers/.pages.

Start with targets ["@summary"], or ["@security"] to triage an untrusted file. Prefer this over unzipping the document or parsing its bytes by hand.

Reports facts ABOUT the document, never its text: it does not extract, render, run macros, follow external references, or open a network connection.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the document. The filename extension selects the format driver, and the container is then verified against it.
viewNoreport: full evidence — value, confidence, path, and cost per target. values: a compact target-to-value map.
levelNoProbe budget and eligible paths. header: identity only. metadata (default). deep: higher-cost paths, needed only when a target's min_level says so.
budgetNoOverride the level's resource limits. Raise after a BUDGET_EXCEEDED error.
targetsNoShort names (slide_count), canonical names (powerpoint.slide_count), or presets: @header, @summary, @security, @structure, @assets, @quality, @format, @all. Defaults to the driver's own set. Call list_targets rather than guessing a name. @summary omits statistics needing a full-file read, so ask for a PDF's page_count explicitly.
min_confidenceNoWeakest evidence a path may offer. Default high; lower it to accept an approximation.
target_confidenceNoPer-target override, e.g. {"slide_count": "exact"} to force the authoritative path.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
schema_versionNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds critical behavioral context beyond those hints: it does not extract, render, run macros, follow external references, or open a network connection; it handles legacy formats; it can hit resource limits (referenced by 'Raise after a BUDGET_EXCEEDED error'). This is exactly the kind of non-obvious behavior an agent needs to trust the tool with untrusted files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: opening scope sentence, format list, usage hint, and a final 'does not' sentence. It is front-loaded with the core purpose and scoping. It earns its sentences; only a minor redundancy exists (the 'Prefer this over...' sentence partially repeats the 'does not' list). Not a 5, but far above the typical terse definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 parameters, nested budget object, output schema present, and complex behavior (format drivers, confidence levels, presets, resource limits), the description provides substantial orientation: preset suggestions, the @summary caveat, the budget-error hint, and a clear non-extraction guarantee. The description does not explain what each target means, but it correctly defers to list_targets. So it's complete enough without being exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 7 parameters well. The description adds important cross-references that the schema alone does not capture: the distinction between @summary and full-read statistics, the warning that @summary omits page_count so ask explicitly, and that list_targets should be used rather than guessing names. This goes beyond a baseline 3 by explaining how the parameters interact with presets and the driver's default set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Read facts about one local PDF, Microsoft Office, or Apple iWork document') plus a precise resource list (page/slide/sheet counts, title/author/dates, encryption, macros, structure, assets, integrity). The scope is explicit: it reports facts ABOUT the document, never extracts text, renders, runs macros, or opens connections. It also names sibling tools (list_targets, probe_batch) and distinguishes from unzipping/byte-parsing by hand. This is clear and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises 'Start with targets ["@summary"], or ["@security"] to triage an untrusted file.' and 'Prefer this over unzipping the document or parsing its bytes by hand.' It also warns about @summary omitting statistics and instructs to 'Call list_targets rather than guessing a name.' This provides direct when-to-use and when-not-to-use guidance, plus pointers to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_batchProbe several documentsA
Read-onlyIdempotent

Probe many local documents in one call — inventory a folder, triage a batch of uploads, or compare a set of files. One engine process handles the whole batch, so this is much cheaper than calling probe once per file.

Takes literal paths; expand any glob yourself first. Returns one entry per path, in the order given, each carrying that file's report or its own error envelope. A file that fails does not affect the others.

Defaults to the compact "values" view because inventory rarely needs per-target evidence; pass view: "report" when it does.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNoreport: full evidence — value, confidence, path, and cost per target. values: a compact target-to-value map.
levelNoProbe budget and eligible paths. header: identity only. metadata (default). deep: higher-cost paths, needed only when a target's min_level says so.
pathsYesDocument paths, at most 64 per call. Literal paths only.
budgetNoOverride the level's resource limits. Raise after a BUDGET_EXCEEDED error.
targetsNoShort names (slide_count), canonical names (powerpoint.slide_count), or presets: @header, @summary, @security, @structure, @assets, @quality, @format, @all. Defaults to the driver's own set. Call list_targets rather than guessing a name. @summary omits statistics needing a full-file read, so ask for a PDF's page_count explicitly.
min_confidenceNoWeakest evidence a path may offer. Default high; lower it to accept an approximation.

Output Schema

ParametersJSON Schema
NameRequiredDescription
reportsYesOne entry per requested path, in order. A per-file failure is confined to it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral detail beyond the annotations: one engine process for the batch, per-path results in order, error envelopes per file, failure isolation, default view and when to switch. Annotations already cover read-only/idempotent/destructive, and the description does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero fluff. The first sentence front-loads purpose and use cases; the second covers path handling and return behavior; the third gives view guidance. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, nested objects, and a full output schema, the description covers the essential usage context: batch behavior, order of results, error isolation, and view default. It does not mention budget override or target selection, but those are adequately documented in the schema. The description is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds valuable semantic guidance: 'Takes literal paths; expand any glob yourself first' clarifies the paths parameter, and the view default rationale ('inventory rarely needs per-target evidence') adds context beyond the schema. It does not elaborate on every parameter, but the key ones receive useful extra meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool probes many local documents in one call, listing concrete use cases (inventory a folder, triage a batch, compare files). It distinguishes itself from the sibling 'probe' by explicitly being the batch counterpart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool ('Probe many local documents in one call'), gives practical scenarios, and contrasts cost efficiency ('much cheaper than calling probe once per file'). It also gives guidance on view selection ('pass view: report when it does').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv0.1.1
    • First observedlist_formats
    • First observedlist_targets
    • First observedprobe
    • First observedprobe_batch

TDQS

A4.7/5.0
Disambiguation5/5

Each tool has a wholly distinct purpose: probe handles a single file, probe_batch handles multiples, list_formats enumerates supported file types, and list_targets enumerates probe targets. There is no overlap or ambiguity between them.

Naming Consistency5/5

All tools follow a consistent lowercase_snake_case convention with clear verb prefixes: probe, probe_batch, list_formats, list_targets. The naming style is uniform and predictable, making the tool set easy to reason about.

Tool Count5/5

With exactly 4 tools, the server is tightly scoped for its purpose of document inspection. Each tool fills a necessary role (single file, batch, format discovery, target discovery) without redundancy or bloat, fitting well within the ideal 3–15 range.

Completeness5/5

The surface covers all core workflows: probing an individual file, probing many files efficiently, discovering supported formats, and enumerating probe targets for a given format. There are no obvious dead ends or missing operations for the server's stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides AI agents with comprehensive document parsing capabilities including PDF text extraction, OCR, HTML-to-markdown conversion, table extraction, and summarization, optimized for agent workflows.
    65
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables Claude and other MCP-compatible agents to process documents, extract structured data, detect PII, and export LLM-ready datasets through natural language tool calls.
    8
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables deterministic visual and structural analysis of PDF and DOCX documents, extracting measurable evidence such as blur, OCR confidence, and image anomalies for auditable forensic workflows.
    1
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/deckflow/deckprobe-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server