Archonics MCP Audit Server
OfficialClick on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Archonics MCP Audit ServerAudit my system prompt for the customer support bot."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Archonics MCP Audit Server
Free-tier context engineering audits for production AI agents, delivered as MCP tools you can call from Claude Desktop, Cursor, Claude Code, or any MCP-compatible client.
What you get: top-3 findings on your system prompts, tool definitions, or context packing, on demand, no account needed.
What it costs: nothing. The free scan is genuinely free. Upgrade paths to the $49 Instant Audit and $750 Full Audit are surfaced in the response footer; they're not paywalls on this tool.
Why this exists
Most production agent failures aren't model failures — they're context engineering failures. Ambiguous instructions, underspecified tools, bloated context, no regression tests on prompt changes. Those problems are spottable by a trained reader. Archonics has trained that reader and published it as an MCP tool so you can get a second opinion on your agent's context without filing a support ticket.
The underlying audit engine applies Archonics Audit Methodology v1.0, the same spec that drives our paid audits.
Related MCP server: ctxray
Tools
audit_system_prompt
Paste a system prompt. Get back the three most important context engineering issues in it, ranked by severity, with specific recommendations.
Covers: role clarity, instruction conflicts, negative space, priority structure when instructions conflict, token efficiency, format specification precision, failure-mode coverage.
audit_tool_definition
Paste a tool/function definition. Get back the three most important issues affecting how reliably the model will call it.
Covers: description quality (the "when to use this tool" question), parameter schema precision, parameter documentation, error response design, discoverability.
audit_context_packing
Paste a representative context payload (or describe it structurally). Get back the three most important efficiency and quality issues.
Covers: content inventory, redundancy across sections, freshness/relevance, ordering, truncation risk, prompt-cache utilization.
Installation
Claude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"archonics-audit": {
"command": "npx",
"args": ["-y", "@archonics/mcp-audit"],
"env": {
"ANTHROPIC_API_KEY": "your-anthropic-api-key-here"
}
}
}
}Cursor
Add to your .cursor/mcp.json:
{
"mcpServers": {
"archonics-audit": {
"command": "npx",
"args": ["-y", "@archonics/mcp-audit"],
"env": {
"ANTHROPIC_API_KEY": "your-anthropic-api-key-here"
}
}
}
}Claude Code
claude mcp add archonics-audit npx -y @archonics/mcp-auditThen set ANTHROPIC_API_KEY in your environment.
Why does it need my Anthropic API key?
The audit engine runs on Claude. You bring your own API key so:
Audit submissions go directly from your machine to Anthropic's API, never through Archonics servers.
Your costs are transparent — a typical audit uses 2,000–4,000 tokens, well under a penny.
There's no "free but actually limited" rate-limit surprise. Your API key, your limits.
If you'd rather not bring your own key, use the $49 Instant Audit at agent.market — we cover the API costs and return a full-methodology audit PDF.
Privacy
Submitted content is processed ephemerally. No prospect content is retained on Archonics infrastructure or used to train any model. The API call pattern is: your client → your Anthropic API key → Anthropic → your client. Archonics servers are not in this path.
Aggregated, anonymized patterns across many audits may inform improvements to the methodology — "18 of 20 audited systems lacked prompt-regression tests" — but specific content never feeds that process.
Details: archonics.ai/privacy
Upgrade paths
If the free scan surfaces issues worth fixing, two paid tiers go deeper:
Instant Audit — $49 USDC via x402. Full methodology applied programmatically to a system you submit. 5-10 page PDF report covering all four dimensions (prompt, tools, context, eval) rather than just three findings in one dimension. Listed at agent.market/archonics.
Full Audit — $750. Human-reviewed audit of a complete agent system. 15-25 page report tuned to your team's context. Contact audits@archonics.ai.
Contact
Questions, feedback, or false-positive reports: audits@archonics.ai
Methodology and full audit examples: archonics.ai
Issues with this MCP server: github.com/archonics/mcp-audit/issues
License
MIT. Use it, fork it, audit yourself.
Available Tools
3 toolsaudit_context_packingA
Analyzes a representative full-context payload and returns the top 3 findings on context efficiency, redundancy, and ordering. Use this when a user is concerned about agent cost, latency, or quality degradation on long conversations. Accepts either a literal dump of what goes into the context window, or a structured description of the context components and their sizes. Findings cover content inventory, redundancy, freshness, ordering, truncation risk, and prompt-cache utilization.
| Name | Required | Description | Default |
|---|---|---|---|
| context_payload | Yes | Either a literal context dump (system prompt + tools + history + retrieved documents as they would appear in an actual API call) OR a structured description like "system prompt: 2400 tokens / tool definitions: 8 tools, ~1800 tokens total / conversation history: last 12 turns, ~6000 tokens / retrieved RAG chunks: top 5, ~3000 tokens." Both formats work; literal dumps produce sharper findings. | |
| context | No | Optional. What kind of agent is this and what is the typical interaction pattern? Single-turn vs. multi-turn, short vs. long conversations, etc. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description discloses that findings cover content inventory, redundancy, freshness, ordering, truncation risk, and prompt-cache utilization. Also notes that literal dumps produce sharper findings. This is sufficient behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, directly stating purpose, usage, and output. No extraneous information. Well structured with key information front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema present, but description explains the coverage of findings (6 aspects). Adequate for a tool that returns a simple text analysis. Could be improved by mentioning number of findings or format, but current is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters have descriptions in the schema. The description adds extra meaning by explaining what kinds of input are acceptable (literal or structured) and notes that 'literal dumps produce sharper findings,' which goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it analyzes context payload for efficiency, redundancy, and ordering. It names specific outputs (top 3 findings) and distinguishes from sibling audit tools by focusing on context packing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this when a user is concerned about agent cost, latency, or quality degradation on long conversations.' It also clarifies acceptable input formats. No explicit exclusion criteria, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_system_promptA
Analyzes a system prompt for context engineering issues and returns the top 3 findings from the Archonics free-tier scan. Use this when a user shares a system prompt from an agent they are building or shipping, especially if they are asking for review, improvement, or debugging help. Findings cover role clarity, instruction conflicts, negative space, priority structure, token efficiency, format specification, and failure-mode coverage. Returns structured JSON with severity-ranked findings. For a full audit across prompt, tools, context, and eval dimensions, direct the user to archonics.ai or the $49 x402 Instant Audit.
| Name | Required | Description | Default |
|---|---|---|---|
| system_prompt | Yes | The complete system prompt text to audit. Paste the full prompt, including any role definitions, instructions, formatting requirements, and examples. Do not redact unless truly necessary; redaction reduces audit quality. | |
| context | No | Optional. Brief description of what the agent is supposed to do and who uses it. One or two sentences. Helps the audit engine assess fit-for-purpose; leaving it blank produces a useful-but-less-targeted audit. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description carries full burden. It discloses return format (structured JSON with severity-ranked findings), scope (top 3 findings), and areas covered. Could mention any limitations like rate limits or caching, but overall transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six sentences, each necessary. Purpose is front-loaded. Could trim minor redundancy but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and moderate complexity, description covers purpose, usage, return format, coverage areas, and alternative options. Thorough for a free-tier scan tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but description adds significant value: for 'system_prompt' it advises to paste full prompt and warns against redaction; for 'context' it explains purpose and impact of leaving blank. Exceeds schema detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb 'Analyzes' and the resource 'system prompt'. It distinguishes from sibling tools by focusing on system prompts, not context packing or tool definitions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: when a user shares a system prompt for review or debugging. Provides an alternative: direct to archonics.ai for full audit. No ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_tool_definitionA
Analyzes a single tool/function definition (name, description, parameter schema) and returns the top 3 findings on tool-call reliability. Use this when a user shares a tool/function definition and asks why the model is calling it wrong, not calling it when expected, or confusing it with other tools. Findings cover description quality, parameter schema precision, parameter documentation, error response design, and discoverability. For auditing an entire tool set together, use the paid tier.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_definition | Yes | The tool definition as it is provided to the model. Accepts JSON schema format (OpenAI-style function calling, Anthropic tool use) or natural-language description. Include the name, description, and parameter schema in full. | |
| context | No | Optional. What agent or system is this tool part of? What other tools does it share a surface with? Helps the audit engine assess overlap and discoverability issues. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It explains that the tool returns top 3 findings covering specific areas (description quality, parameter schema precision, etc.) and accepts various formats. While it does not mention potential errors or limitations, for a read-only analysis tool the description provides sufficient behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus a brief list of coverage areas. It is front-loaded with the main action, contains no filler, and every sentence adds useful information. The structure is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no output schema), the description is complete. It explains what the tool does, when to use it, what it returns (top 3 findings), and the aspects it covers. It also provides guidance on alternatives, making it self-contained for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value beyond the schema by specifying that the tool_definition parameter accepts JSON schema or natural-language format, and that context is optional for assessing overlap and discoverability. This extra information enhances parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes a single tool/function definition and returns top 3 findings on tool-call reliability. It distinguishes itself from sibling tools by focusing on individual definitions, and mentions a paid tier for auditing entire tool sets, implying this tool is for single definitions only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'when a user shares a tool/function definition and asks why the model is calling it wrong, not calling it when expected, or confusing it with other tools.' It also provides an alternative: 'For auditing an entire tool set together, use the paid tier.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v0.1.5- First observed
audit_context_packing - First observed
audit_system_prompt - First observed
audit_tool_definition
TDQS
Each tool targets a distinct audit dimension: context packing, system prompt, and tool definition. There is no overlap in purpose, making it clear which tool to use for a given issue.
All tool names follow a consistent 'audit_<target>' pattern using snake_case, which is predictable and clearly indicates the function of each tool.
Three tools is appropriate for a specialized audit server, covering the core areas of context, system prompt, and tool definition without unnecessary bloat.
The server covers three critical audit aspects, but lacks a tool for full tool-set coherence or evaluation, which is noted as a paid-tier feature. Minor gap in free tier completeness.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Free public MCP for AI agents — 193 tools, 44 workflows. No API key.
Free MCP tools: the only MCP linter, health checks, cost estimation, and trust evaluation.
MCP-first toolbox for agents: KV storage, auth, queue, and utility tools. Free in early access.
AI Visibility and Content Intelligence tools for Claude and MCP-compatible agents.
Related MCP Servers
- AlicenseAqualityFmaintenanceScore your agent's governance (0-100), lint MCP tool definitions, and estimate costs across all major models. Free diagnostic tools with no API key needed. Expert skill files on governance, economics, and system architecture available with free tier.81MIT
- AlicenseNot gradedqualityCmaintenanceContext intelligence for AI coding sessions. 7 MCP tools to score, compare, compress, build, and scan prompts across 9 AI tools. Rule-based, <5ms/prompt, all analysis runs locally.46MIT
- AlicenseAqualityBmaintenanceMCP tools for video transcoding, document conversion, and multi-step pipelines — callable by any AI agent.121371MIT
- FlicenseNot gradedqualityAmaintenanceProvides real-time monitoring of AI agents, context, usage limits, workflows, files, Git, tests, builds, errors, secrets, and model-economy advice for tools like Claude Code, Codex, and Cursor, with 30 MCP tools for comprehensive observability.1-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/archonics/mcp-audit'
If you have feedback or need assistance with the MCP directory API, please join our Discord server