PromptShield MCP
This server provides MCP tools for AI agents and firewalls to scan text for PII, secrets, and harmful content, redact sensitive information, and get policy recommendations. It offers:
promptshield.scan_text: Scan any text string (prompts, outputs, logs) for PII (e.g., emails), secrets (e.g., API keys), and harmful content, returning findings with risk levels, confidence scores, and character spans.promptshield.scan_messages: Scan chat-style messages while preserving roles and indexes, useful for inspecting conversations before sending to models or tools.promptshield.redact_text: Scan and return a redacted version of text with sensitive spans replaced (e.g.,[PII],[SECRET]), along with a list of findings.promptshield.evaluate_policy: Map findings to an action recommendation (allow,warn,redact,block) for consistent enforcement decisions.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PromptShield MCPScan this text for PII: 'email@test.com'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Zero Harm AI MCP
Zero Harm AI MCP is a Model Context Protocol server that lets AI agents and runtime firewalls call Zero Harm AI safety checks for text, chat messages, prompts, tool inputs, and generated outputs.
The server is a thin adapter over zero-harm-ai-detectors. It should not duplicate detector logic from the detector package or from the Zero Harm AI GitHub Action.
Goals
Expose PII, secret, and harmful-content detection through MCP tools.
Return structured findings that agents and firewalls can enforce.
Support local/self-hosted operation for sensitive data.
Keep logs privacy-safe by default.
Provide stable tool contracts that can be used by coding agents, chat agents, and firewall.
Related MCP server: AIM-Guard-MCP
Non-Goals
Reimplementing
zero-harm-ai-detectors.Acting as a hosted service by default.
Making policy enforcement decisions that belong to a firewall or calling agent.
Replacing the Zero Harm AI GitHub Action.
Relationship To Other Projects
zero-harm-ai-detectors
Shared detector engine for PII, secrets, and harmful content.
zero-harm-ai-gh-action
GitHub Action and CI-oriented scanner for pull requests.
zero-harm-ai-mcp
MCP server adapter that exposes detector functionality to AI agents.
zero-harm-ai-firewall (future)
Runtime enforcement layer. It can call zero-harm-ai-mcp or use
zero-harm-ai-detectors directly.Installation
Install from PyPI:
pip install zero-harm-ai-mcpFor local development:
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
ruff check .MCP Client Configuration
{
"mcpServers": {
"zero-harm-ai": {
"command": "zero-harm-ai-mcp",
"args": []
}
}
}MCP Tools
zero_harm.scan_text
Scan one text string for PII, secrets, and harmful content.
Use this for prompt inputs, generated outputs, tool arguments, logs, and arbitrary text.
zero_harm.scan_messages
Scan chat-style messages while preserving message roles and indexes.
Use this when an agent wants to inspect a conversation before sending it to a model or tool.
zero_harm.redact_text
Return a redacted version of text plus findings.
Use this when the caller wants to continue safely after removing sensitive spans.
zero_harm.evaluate_policy
Map detector findings to an action recommendation.
Use this when a caller wants a normalized decision such as allow, warn, redact, or block.
Working Examples
These examples are generated from the current local server implementation.
zero_harm.scan_text
Input:
{
"text": "Contact alice@example.com before sharing the token.",
"targets": ["pii", "secret", "harmful"],
"redact": true
}Output:
{
"schema_version": "1.0.0",
"risk_level": "medium",
"recommended_action": "redact",
"categories": [
"pii"
],
"summary": {
"total_findings": 1,
"pii": 1,
"secret": 0,
"harmful": 0
},
"findings": [
{
"type": "email",
"category": "pii",
"severity": "medium",
"confidence": 0.99,
"span": {
"start": 8,
"end": 25
},
"redacted": "[PII]",
"message_index": null,
"message_role": null,
"evidence_available": false
}
],
"redacted_text": "Contact [PII] before sharing the token."
}zero_harm.scan_messages
Input:
{
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "My email is alice@example.com."
}
],
"targets": ["pii", "secret", "harmful"],
"redact": true
}Output:
{
"schema_version": "1.0.0",
"risk_level": "medium",
"recommended_action": "redact",
"categories": [
"pii"
],
"summary": {
"total_findings": 1,
"pii": 1,
"secret": 0,
"harmful": 0
},
"findings": [
{
"type": "email",
"category": "pii",
"severity": "medium",
"confidence": 0.99,
"span": {
"start": 12,
"end": 29
},
"redacted": "[PII]",
"message_index": 1,
"message_role": "user",
"evidence_available": false
}
],
"redacted_text": "[{\"role\": \"system\", \"content\": \"You are a helpful assistant.\"}, {\"role\": \"user\", \"content\": \"My email is [PII].\"}]"
}zero_harm.redact_text
Input:
{
"text": "aws_access_key_id = AKIAIOSFODNN7EXAMPLE",
"targets": ["pii", "secret", "harmful"]
}Output:
{
"schema_version": "1.0.0",
"risk_level": "high",
"recommended_action": "block",
"categories": [
"secret"
],
"summary": {
"total_findings": 1,
"pii": 0,
"secret": 1,
"harmful": 0
},
"findings": [
{
"type": "api_key",
"category": "secret",
"severity": "high",
"confidence": 0.95,
"span": {
"start": 20,
"end": 40
},
"redacted": "[SECRET]",
"message_index": null,
"message_role": null,
"evidence_available": false
}
],
"redacted_text": "aws_access_key_id = [SECRET]"
}zero_harm.evaluate_policy
Input:
{
"text": "Contact alice@example.com before sharing the token.",
"targets": ["pii", "secret", "harmful"],
"redact": false
}Output:
{
"schema_version": "1.0.0",
"risk_level": "medium",
"recommended_action": "warn",
"categories": [
"pii"
],
"summary": {
"total_findings": 1,
"pii": 1,
"secret": 0,
"harmful": 0
}
}Privacy Requirements
Do not log raw input text by default.
Do not log detected secret values by default.
Include a config option for audit logs that stores only counts, categories, severities, and request metadata.
Avoid sending data to external services unless explicitly configured.
Keep the default transport local-first.
Development
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
ruff check .Release
Build and validate distribution artifacts:
python -m build
twine check dist/*See RELEASE.md for the full PyPI release flow.
Available Tools
4 toolspromptshield.evaluate_policyC
Return a normalized action recommendation for the supplied text.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and description only states output behavior. Does not disclose side effects, authorization needs, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but severely under-specifies the tool's functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 0% schema coverage, no output schema, and no annotations, the description fails to provide necessary context for proper invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description does not explain the required 'payload' object or its expected structure. The schema allows any properties, so the agent has no guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it returns a normalized action recommendation, but 'normalized action' and 'policy' are vague. It distinguishes from sibling scanning tools but lacks specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings like scan_text or redact_text. No context on prerequisites or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
promptshield.redact_textC
Scan and redact sensitive spans from text.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only states the action, omitting details on what constitutes sensitive spans, redaction format, side effects (e.g., mutability), permission requirements, or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and front-loaded. However, it is too minimal to be maximally effective, lacking necessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one complex parameter, no schema descriptions, no output schema, and no annotations, the description is severely incomplete. It fails to explain how to use the tool, what inputs are expected, or what outputs occur.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter 'payload' is an object with additionalProperties true and no schema description. The description adds no meaning beyond 'from text', failing to specify expected structure or how to provide text for redaction.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it scans and redacts sensitive spans from text, which is a clear verb+resource combination. It distinguishes from siblings: evaluate_policy, scan_messages, and scan_text do not mention redaction. However, 'sensitive spans' is somewhat vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus its siblings. The description provides no context for when redaction is appropriate or when to prefer alternative tools like scan_text for scanning only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
promptshield.scan_messagesC
Scan chat messages while preserving role and message indexes.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It states that role and message indexes are preserved, which is a positive behavioral trait. However, it does not disclose whether the tool modifies data, requires authentication, or has side effects, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise, but it sacrifices helpful detail. It is front-loaded with the key point, but may be too terse to be useful without additional context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one free-form parameter, no output schema, no annotations, and multiple siblings, the description is far from complete. It lacks details on the return format, how the payload should be structured, and how it differs from promptshield.scan_text.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero description coverage (0%), and the description adds no meaning to the 'payload' parameter. The payload is a free-form object with no defined properties, and the description fails to explain its expected structure or content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies 'scan chat messages' and mentions preserving role and message indexes, which clearly indicates the resource (chat messages) and a key behavior. However, the verb 'scan' is somewhat ambiguous (could imply reading, checking, or analyzing), and it doesn't fully distinguish from sibling tools like 'scan_text'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, exclusions, or comparison to sibling tools like promptshield.evaluate_policy or promptshield.scan_text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
promptshield.scan_textC
Scan text for PII, secrets, and harmful content.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must disclose behavior. It mentions what is scanned for but omits details on return value, side effects (read-only or destructive), or whether it modifies the input. The description is too minimal to fully inform an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence, but it sacrifices necessary detail for brevity. While concise, it is under-specified for the complexity of the tool (nested object parameter, no schema description).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations, output schema, and 0% schema parameter coverage, the description is extremely incomplete. It fails to explain the parameter, return format, or how it compares to siblings. The tool is not adequately specified for reliable agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explain the 'payload' parameter. It is an object with additionalProperties: true, but no guidance on expected structure (e.g., keys, format of text). Agents cannot infer how to correctly invoke the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans text for PII, secrets, and harmful content. The verb 'scan' and target 'text' are explicit, and it distinguishes from siblings like 'redact_text' (modification) and 'evaluate_policy' (policy evaluation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like promptshield.scan_messages or promptshield.redact_text. No context about prerequisites or typical use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v0.1.0- First observed
promptshield.evaluate_policy - First observed
promptshield.redact_text - First observed
promptshield.scan_messages - First observed
promptshield.scan_text
TDQS
Each tool targets a distinct operation (evaluate policy, redact, scan text, scan messages) with clear descriptions, making them easily distinguishable.
All tools follow a consistent 'promptshield.<verb>_<object>' pattern with lowercase underscores, using verbs like evaluate, redact, and scan.
Four tools is appropriate for a content safety server, covering key tasks without being too many or too few.
The set covers evaluation, scanning, and redaction for both general text and chat messages, providing a complete surface for common content security needs.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Email safety MCP server. Detects phishing, prompt injection, CEO fraud for AI agents.
A Model Context Protocol server for Wix AI tools
Security firewall for AI agents — scans MCP calls for injection, secrets, and risks.
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA minimal Model Context Protocol server that provides a safety guardrail tool to check if provided context is free from code injection or harmful content.-
- AlicenseCqualityCmaintenanceA Model Context Protocol (MCP) server that provides AI-powered security analysis and safety instruction tools. This server helps protect AI agents by providing security guidelines, content analysis, and cautionary instructions when interacting with various MCPs and external services.65321ISC
- AlicenseAqualityFmaintenanceUnified MCP safety server that detects prompt injection (75 patterns), scans LLM outputs for leaked secrets/PII, enforces API cost budgets, and creates signed audit trails. Zero ML dependencies, pure Python.171MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for AI agent security guardrails. Provides input validation, prompt injection detection, PII redaction, output filtering, policy enforcement, rate limiting, and comprehensive audit logging.761MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Zero-Harm-AI/zero-harm-ai-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server