DeckProbe MCP Server
OfficialClick on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DeckProbe MCP ServerWhat's the slide count and security summary for deck.pptx?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DeckProbe MCP Server
Let an agent ask what's inside a PDF, Office, or iWork file — without opening it.
Install · Tools · Configuration · Security · How it works · DeckProbe
An MCP server that exposes
DeckProbe — ffprobe for documents — as
four typed tools. Ask for page counts, slide counts, metadata, encryption and
macro signals, structure, or integrity, and get back bounded, deterministic JSON
with confidence, evidence, and measured I/O cost.
Nothing is rendered, no macro runs, no external reference is followed, and no network connection is opened. It is safe to point at untrusted files.
// probe { "path": "deck.pptx", "targets": ["slide_count"], "view": "values" }
{
"schema_version": 2,
"status": "ok",
"driver": { "id": "powerpoint", "profile": "pptx" },
"values": { "powerpoint.slide_count": 31 },
"view": "values"
}Install
Nothing to install ahead of time — npx fetches the server and the engine
together.
Claude Code
claude mcp add deckprobe -- npx -y @deckflow/deckprobe-mcpClaude Desktop, Cursor, VS Code, Zed, and anything else reading mcpServers
{
"mcpServers": {
"deckprobe": {
"command": "npx",
"args": ["-y", "@deckflow/deckprobe-mcp"]
}
}
}For a pinned install, npm install -g @deckflow/deckprobe-mcp and use
deckprobe-mcp as the command.
Requires Node.js 20 or newer. The engine binary arrives as a per-platform
optional dependency for macOS, Linux (glibc and musl), and Windows on x86-64 and
ARM64; anywhere else the server falls back to the same engine compiled to
WebAssembly, so npx works wherever Node does.
Related MCP server: flexorch-mcp
Tools
Tool | Use it for |
| Everything about one document |
| Inventory or triage many documents in one call |
| Which formats are supported, and where support stops |
| The exact target names a format offers |
There is also one resource, deckprobe://schema, carrying the report JSON
Schema bundled with the running engine.
probe
{
"path": "reports/q3.pptx",
"targets": ["@summary", "@security"], // presets, short names, or canonical names
"level": "metadata", // header | metadata | deep
"min_confidence": "high", // low | medium | high | exact
"target_confidence": { "slide_count": "exact" },
"view": "report", // report | values
"budget": { "max_physical_bytes": 8388608, "timeout_ms": 1000 }
}targets accepts short names (slide_count), canonical names
(powerpoint.slide_count), and presets:
Preset | Expands to |
| Container identity only — format, size, extension match, encryption flag |
| Identity, common metadata, and primary structure |
| Encryption, macros, signatures, external references, active content |
| Format-owned counts, names, and dimensions |
| Images, media, previews, fonts, embedded objects |
| Integrity, repair, extension match, conformance |
| Every format-specific target at the active level |
| Everything available at the active level |
@summary deliberately omits statistics that need a full-file read. A PDF's
page_count is the notable case — ask for it explicitly.
probe_batch
{ "paths": ["a.pdf", "b.pptx", "c.xlsx"], "targets": ["@security"] }One engine process handles the whole batch. Results come back in input order,
each with its own report or its own error, so one bad file never spoils the run.
Defaults to the compact values view. Literal paths only — expand globs
yourself.
list_formats and list_targets
list_targets takes a format (pdf, docx, xlsx, pptx, doc, xls,
ppt, key, numbers, pages) and returns each target's aliases,
description, value type, minimum level, cost class, and selector membership.
Pass detail: "full" for the engine's complete report, including per-target
JSON Schema fragments and expanded selector lists.
Both are cached for the lifetime of the server process.
Reading a report
The tool result is the engine's own schema-v2 envelope, unmodified. Two things are worth knowing before consuming it:
status: "partial"is not a failure. It means at least one requested target could not be resolved at the requested confidence. It is named inexecution.unresolved_targets, and every other result still stands.confidence_scoreis a fixed constant per label (0.4,0.7,0.95,1.0), not a calibrated probability.0.95does not mean the value is right 95% of the time.
Only results with status resolved or estimated carry a value. unknown is
common and usually means the document simply does not record that fact.
A failing call returns isError with the engine's error envelope plus one line
saying what to do about it. Failures the server itself raises before the engine
runs — a missing path, a directory, a path outside the allow-list, an exceeded
deadline — use the same envelope shape with an MCP_-prefixed code and
origin: "mcp-server".
Configuration
Every setting is an environment variable, set in your client's MCP config. All are optional.
Variable | Default | Meaning |
| – | Engine binary to use instead of the bundled one |
| unrestricted | Allowed directories, separated like |
|
| Hard per-call deadline on an engine process |
|
| Concurrent engine processes |
|
| Paths accepted by one |
{
"deckprobe": {
"command": "npx",
"args": ["-y", "@deckflow/deckprobe-mcp"],
"env": { "DECKPROBE_MCP_ROOTS": "/Users/me/Documents:/Users/me/Downloads" }
}
}Security
DeckProbe is built for untrusted input: bounded parsing, no renderer, no macro interpreter, no external-reference resolution, and no network access. This server adds two things on top.
Process isolation and a hard deadline. Each probe runs in its own short-lived process, killed if it outruns
DECKPROBE_MCP_TIMEOUT_MS.An optional read allow-list.
DECKPROBE_MCP_ROOTSpins the reachable tree; paths are symlink-resolved before the check, so a link cannot step around it. The default is unrestricted, matching the CLI the user could run themselves — set it for shared or automated deployments.
Reports describe a document (metadata, counts, signals) rather than reproducing its contents. Note that report values such as a document title are still attacker-controlled strings: the server passes them through as JSON data and never interpolates them into instructions, and a consumer should treat them the same way.
Report a vulnerability privately as described in SECURITY.md.
How it works
MCP client
│ JSON-RPC over stdio
▼
deckprobe-mcp ── validates arguments, resolves the path, maps the result
│ argv + stdout (one process per probe, or one --jsonl process per batch)
▼
DeckProbe engine ── plans the cheapest paths that answer the requestThe server spawns the native DeckProbe CLI rather than calling the WebAssembly build. The CLI reads only the byte ranges a probe plan needs, where the WebAssembly path holds the whole file in memory, and a separate OS process both isolates untrusted parsing and can be killed outright. The engine is chosen in this order:
DECKPROBE_MCP_BINthe binary that ships with this package's
@deckflow/deckprobedependencydeckprobeonPATHthe bundled WebAssembly engine
The resolved engine is logged to stderr at startup. stdout belongs to the MCP transport and carries nothing else.
MCP server or agent skill?
DeckProbe also ships an Agent Skill that teaches a shell-capable agent to use the CLI directly. Both teach the same vocabulary. Use the skill when the agent has a shell and you want the CLI's full surface; use this server when it does not, or when you want typed arguments validated before the engine ever runs.
Development
npm install
npm test # typecheck, lint, build, and the full suite
npm run test:watchContributions are welcome — see CONTRIBUTING.md. The design rationale, including the alternatives that were rejected, is in docs/rfc.md.
License
MIT. See LICENSE.
Available Tools
4 toolslist_formatsList supported formatsARead-onlyIdempotent
List the document formats DeckProbe can inspect: drivers, the extensions each one handles, and where support stops. Call this when you are unsure whether a file type is supported at all.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, non-destructive behavior. The description adds valuable context about the exact output content (drivers, extensions, support limitations), which goes beyond the annotations. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, with the primary purpose and scope front-loaded. It avoids redundancy and every sentence earns its place—the first states what it lists, the second when to call it. No unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless list tool with no output schema, the description fully explains what the tool returns (drivers, extensions, and support limits). There is nothing missing for an agent to correctly invoke it and interpret the result. Complexity is low, so this is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is fully covered (100% by default). There is nothing to add about parameters; the description doesn't need to explain any. The baseline for zero parameters is 4, and no additional info is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('document formats DeckProbe can inspect'), and specifies the content (drivers, extensions, and where support stops). It effectively distinguishes itself from sibling tools like probe and list_targets by focusing on format capabilities rather than probing or target listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Call this when you are unsure whether a file type is supported at all.' It doesn't mention alternatives, but the trigger condition is clear and implies that if you have a specific file, you would use probe instead. It could be improved by naming the sibling tools explicitly, but the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_targetsList targets for a formatARead-onlyIdempotent
List every target a format supports, with its aliases, value type, minimum probe level, cost class, and which @selectors include it. Call this before naming a target you have not already seen — probe rejects an unknown one rather than guessing.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | compact (default): target names, aliases, descriptions, levels, and selector membership. full: the complete report, including each target's JSON Schema fragment and every selector's expanded member list. | |
| format | Yes | A format profile from list_formats, such as "pdf", "docx", "xlsx", "pptx", "doc", "xls", "ppt", "key", "numbers", or "pages". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, covering safety. The description adds behavioral context by spelling out returned fields and the probe rejection behavior, which matters for planning calls. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences front-load the core purpose and output contents, then give one practical usage rule. No filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich annotation set, fully documented parameters, and a description that names the output fields and the prerequisite call, an agent has enough to invoke the tool correctly even with no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already explains format and detail fully, including enumerations and examples. The description adds no new parameter-level semantics beyond naming the report contents, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (list) and resource (targets for a given format), enumerates exactly what is included (aliases, value type, minimum probe level, cost class, selector membership), and contrasts with probe by stating probe rejects unknown targets. This distinguishes it from sibling tools list_formats and probe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs when to call: before naming a target not already seen, because probe rejects unknowns rather than guessing. It also ties the format parameter to list_formats, implying the prerequisite and differentiating from list_formats.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probeProbe a documentARead-onlyIdempotent
Read facts about one local PDF, Microsoft Office, or Apple iWork document without opening or rendering it: page, slide, and sheet counts; title, author, and dates; encryption, macro, signature, and active-content signals; structure; embedded assets; and integrity. Handles .pdf, .docx/.xlsx/.pptx, legacy .doc/.xls/.ppt, and modern .key/.numbers/.pages.
Start with targets ["@summary"], or ["@security"] to triage an untrusted file. Prefer this over unzipping the document or parsing its bytes by hand.
Reports facts ABOUT the document, never its text: it does not extract, render, run macros, follow external references, or open a network connection.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path to the document. The filename extension selects the format driver, and the container is then verified against it. | |
| view | No | report: full evidence — value, confidence, path, and cost per target. values: a compact target-to-value map. | |
| level | No | Probe budget and eligible paths. header: identity only. metadata (default). deep: higher-cost paths, needed only when a target's min_level says so. | |
| budget | No | Override the level's resource limits. Raise after a BUDGET_EXCEEDED error. | |
| targets | No | Short names (slide_count), canonical names (powerpoint.slide_count), or presets: @header, @summary, @security, @structure, @assets, @quality, @format, @all. Defaults to the driver's own set. Call list_targets rather than guessing a name. @summary omits statistics needing a full-file read, so ask for a PDF's page_count explicitly. | |
| min_confidence | No | Weakest evidence a path may offer. Default high; lower it to accept an approximation. | |
| target_confidence | No | Per-target override, e.g. {"slide_count": "exact"} to force the authoritative path. |
Output Schema
| Name | Required | Description |
|---|---|---|
| status | No | |
| schema_version | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds critical behavioral context beyond those hints: it does not extract, render, run macros, follow external references, or open a network connection; it handles legacy formats; it can hit resource limits (referenced by 'Raise after a BUDGET_EXCEEDED error'). This is exactly the kind of non-obvious behavior an agent needs to trust the tool with untrusted files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: opening scope sentence, format list, usage hint, and a final 'does not' sentence. It is front-loaded with the core purpose and scoping. It earns its sentences; only a minor redundancy exists (the 'Prefer this over...' sentence partially repeats the 'does not' list). Not a 5, but far above the typical terse definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, nested budget object, output schema present, and complex behavior (format drivers, confidence levels, presets, resource limits), the description provides substantial orientation: preset suggestions, the @summary caveat, the budget-error hint, and a clear non-extraction guarantee. The description does not explain what each target means, but it correctly defers to list_targets. So it's complete enough without being exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters well. The description adds important cross-references that the schema alone does not capture: the distinction between @summary and full-read statistics, the warning that @summary omits page_count so ask explicitly, and that list_targets should be used rather than guessing names. This goes beyond a baseline 3 by explaining how the parameters interact with presets and the driver's default set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Read facts about one local PDF, Microsoft Office, or Apple iWork document') plus a precise resource list (page/slide/sheet counts, title/author/dates, encryption, macros, structure, assets, integrity). The scope is explicit: it reports facts ABOUT the document, never extracts text, renders, runs macros, or opens connections. It also names sibling tools (list_targets, probe_batch) and distinguishes from unzipping/byte-parsing by hand. This is clear and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises 'Start with targets ["@summary"], or ["@security"] to triage an untrusted file.' and 'Prefer this over unzipping the document or parsing its bytes by hand.' It also warns about @summary omitting statistics and instructs to 'Call list_targets rather than guessing a name.' This provides direct when-to-use and when-not-to-use guidance, plus pointers to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probe_batchProbe several documentsARead-onlyIdempotent
Probe many local documents in one call — inventory a folder, triage a batch of uploads, or compare a set of files. One engine process handles the whole batch, so this is much cheaper than calling probe once per file.
Takes literal paths; expand any glob yourself first. Returns one entry per path, in the order given, each carrying that file's report or its own error envelope. A file that fails does not affect the others.
Defaults to the compact "values" view because inventory rarely needs per-target evidence; pass view: "report" when it does.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | report: full evidence — value, confidence, path, and cost per target. values: a compact target-to-value map. | |
| level | No | Probe budget and eligible paths. header: identity only. metadata (default). deep: higher-cost paths, needed only when a target's min_level says so. | |
| paths | Yes | Document paths, at most 64 per call. Literal paths only. | |
| budget | No | Override the level's resource limits. Raise after a BUDGET_EXCEEDED error. | |
| targets | No | Short names (slide_count), canonical names (powerpoint.slide_count), or presets: @header, @summary, @security, @structure, @assets, @quality, @format, @all. Defaults to the driver's own set. Call list_targets rather than guessing a name. @summary omits statistics needing a full-file read, so ask for a PDF's page_count explicitly. | |
| min_confidence | No | Weakest evidence a path may offer. Default high; lower it to accept an approximation. |
Output Schema
| Name | Required | Description |
|---|---|---|
| reports | Yes | One entry per requested path, in order. A per-file failure is confined to it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral detail beyond the annotations: one engine process for the batch, per-path results in order, error envelopes per file, failure isolation, default view and when to switch. Annotations already cover read-only/idempotent/destructive, and the description does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero fluff. The first sentence front-loads purpose and use cases; the second covers path handling and return behavior; the third gives view guidance. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters, nested objects, and a full output schema, the description covers the essential usage context: batch behavior, order of results, error isolation, and view default. It does not mention budget override or target selection, but those are adequately documented in the schema. The description is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds valuable semantic guidance: 'Takes literal paths; expand any glob yourself first' clarifies the paths parameter, and the view default rationale ('inventory rarely needs per-target evidence') adds context beyond the schema. It does not elaborate on every parameter, but the key ones receive useful extra meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool probes many local documents in one call, listing concrete use cases (inventory a folder, triage a batch, compare files). It distinguishes itself from the sibling 'probe' by explicitly being the batch counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use this tool ('Probe many local documents in one call'), gives practical scenarios, and contrasts cost efficiency ('much cheaper than calling probe once per file'). It also gives guidance on view selection ('pass view: report when it does').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v0.1.1- First observed
list_formats - First observed
list_targets - First observed
probe - First observed
probe_batch
TDQS
Each tool has a wholly distinct purpose: probe handles a single file, probe_batch handles multiples, list_formats enumerates supported file types, and list_targets enumerates probe targets. There is no overlap or ambiguity between them.
All tools follow a consistent lowercase_snake_case convention with clear verb prefixes: probe, probe_batch, list_formats, list_targets. The naming style is uniform and predictable, making the tool set easy to reason about.
With exactly 4 tools, the server is tightly scoped for its purpose of document inspection. Each tool fills a necessary role (single file, batch, format discovery, target discovery) without redundancy or bloat, fitting well within the ideal 3–15 range.
The surface covers all core workflows: probing an individual file, probing many files efficiently, discovering supported formats, and enumerating probe targets for a given format. There are no obvious dead ends or missing operations for the server's stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Agent-native document parsing: PDF, scans and FR/EU invoices to structured JSON or Markdown.
Check a document for hidden text before your agent reads it. PDF, Office, RTF, HTML.
1Verified OCR with per-value coordinates, plus a workspace agents can file documents into and query.
Agent microtools: secret scanning, preflight, receipts, JSON checks and Zero Point analysis.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides AI agents with comprehensive document parsing capabilities including PDF text extraction, OCR, HTML-to-markdown conversion, table extraction, and summarization, optimized for agent workflows.65MIT

flexorch-mcpofficial
AlicenseAqualityAmaintenanceEnables Claude and other MCP-compatible agents to process documents, extract structured data, detect PII, and export LLM-ready datasets through natural language tool calls.81MIT- FlicenseNot gradedqualityBmaintenanceEnables deterministic visual and structural analysis of PDF and DOCX documents, extracting measurable evidence such as blur, OCR confidence, and image anomalies for auditable forensic workflows.1-
- AlicenseAqualityCmaintenanceEnables AI agents to process and inspect PDFs by detecting document type, extracting text, and generating Markdown with layout information.71MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/deckflow/deckprobe-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server