Token Guardian MCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Token Guardian MCPshow me my token usage for the last 7 days, broken down by agent and model"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Token Guardian MCP
Token Guardian is a local, read-only MCP server for finding Claude Code and Codex token waste without making either client dumber.
It does two jobs:
Reports which agents, models, and sessions are processing the most tokens.
Recommends the cheapest safe model and effort level for a specific task.
It also includes a dry-run route-task command for testing prompt-first routing before the MCP is registered anywhere.
Hard safety boundary
Token Guardian cannot change model settings, effort levels, context limits, compaction, MCP registration, or running sessions. It has no model credentials and makes no paid API calls.
The usage tool calls ccusage with --offline and --no-cost. Codex titles are optional metadata read through sqlite3 -readonly. Every MCP tool is marked read-only, non-destructive, idempotent, and closed-world.
Related MCP server: token-pilot
Tools
token_guardian_usage_snapshot
Reads one through thirty days of local usage and returns:
Totals by agent and model
Cache-read, output, and frontier-model shares
Largest sessions
Exact duplicate Codex work titles
Evidence-backed quick wins
Processed tokens are a workload diagnostic. They are not the same as a subscription meter, especially when cached input dominates.
token_guardian_recommend_route
Accepts the client, task, risk, and current context size. It returns a model, effort level, reasons, and an optional frontier validator.
The policy fails closed. Security, architecture, production, destructive, high-risk, critical, or ambiguous work stays on a frontier model. Cheap models handle bounded mechanical work. Balanced models handle normal coding and debugging, with a frontier review when needed.
Test prompt-first routing
Build the project, then run the router from any folder:
node ./dist/route-cli.js --client codex --prompt "Implement a bounded TypeScript parser with tests." --cwd .The working folder does not need to be a repository. The router primarily uses the prompt. It checks up to 256 names in the current folder for optional project markers, without opening file contents or scanning subfolders.
The command only prints a recommendation and a session-scoped launch command. It does not run that command, change configuration, register the MCP, or touch an active session. Vague continuation prompts such as yeah, test that stay on the frontier route because their real context is missing.
Local requirements
Node.js 24 or newer
ccusageon PATHsqlite3on PATH for optional Codex thread titles
Build and test
npm install
npm test
npm run typecheck
npm run buildRun the built stdio server:
node dist/index.jsInspect it without registering it in a live client:
npx @modelcontextprotocol/inspector node dist/index.jsRegistration status
This first build is intentionally not registered with Claude, Codex, or the shared Kitsune gateway. Registration and any client restart are a separate change after the isolated server is accepted.
License
MIT. See LICENSE.
Available Tools
2 toolstoken_guardian_recommend_routeRecommend Model RouteARead-onlyIdempotent
Recommend a conservative Claude, Codex, or local model and effort for one task. Returns advice only and never switches the active client.
| Name | Required | Description | Default |
|---|---|---|---|
| risk | No | normal | |
| client | Yes | Client that will perform the task. | |
| task_kind | No | Explicit task class. Omit only when genuinely unknown. | |
| task_summary | Yes | Short description of the actual work. | |
| current_model | No | ||
| context_tokens | No | ||
| current_effort | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| tier | Yes | |
| model | Yes | |
| client | Yes | |
| effort | Yes | |
| actions | Yes | |
| reasons | Yes | |
| validator | No | |
| confidence | Yes | |
| advisoryOnly | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable behavioral context: 'Returns advice only and never switches the active client', confirming the tool's side-effect-free nature. No contradiction with annotations. The description could mention error handling or latency but is sufficient given the strong annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the key action, and contains no filler. Every word adds value: 'conservative' qualifies the recommendation, 'one task' sets scope, and 'never switches' clarifies behavior. Ideal conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, output schema exists, good annotations), the description covers the core workflow but leaves out context like the meaning of 'conservative', 'effort', or how the recommendation relates to the input parameters. The output schema might fill gaps, but the description could be more helpful for understanding the tool's role in a multi-step workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43% (3 of 7 parameters have descriptions). The description adds no parameter-specific information; it only states the overall purpose. Parameters like 'risk', 'current_effort', and 'context_tokens' remain undocumented, and the description does not compensate for the low coverage. The agent must rely on the incomplete schema, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'recommend', the resource ('model and effort'), and the scope ('for one task'). It also distinguishes from the sibling tool 'token_guardian_usage_snapshot' by emphasizing that it returns advice only and does not change state, making the purpose specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus the sibling 'token_guardian_usage_snapshot'. It states that it returns advice only, implying safe usage, but does not mention when to choose recommendations over usage snapshots or any prerequisites or exclusions. This leaves the agent without clear direction on invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
token_guardian_usage_snapshotToken Usage SnapshotARead-onlyIdempotent
Read local Claude Code and Codex usage, identify large contexts and expensive routing, and return evidence-backed quick wins. Never changes settings or sessions.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | ||
| agent | No | all | |
| top_sessions | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| caveat | Yes | |
| totals | Yes | |
| window | Yes | |
| byAgent | Yes | |
| byModel | Yes | |
| metrics | Yes | |
| quickWins | Yes | |
| topSessions | Yes | |
| duplicateCodexWork | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint as safe. The description reinforces this by stating 'Never changes settings or sessions.' It adds value by describing what the tool does with the data (identify, return quick wins), which goes beyond the annotations' safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences that pack in the tool's purpose, outcomes, and safety guarantee. Every word serves a purpose with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a read-only snapshot tool with an output schema and sensible defaults, but it fails to mention the tunable parameters (days, agent, top_sessions) that control the snapshot scope. The sibling tool is not contrasted, leaving the agent to infer the relationship.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not mention the three parameters (days, agent, top_sessions) or their roles. The parameter names are somewhat self-explanatory, but the description adds no additional meaning beyond the schema's type and default constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads usage data, identifies large contexts and expensive routing, and returns quick wins. It uses a specific verb ('Read') and resource ('local Claude Code and Codex usage'), and distinguishes from the sibling tool 'token_guardian_recommend_route' by focusing on analysis rather than recommendations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states it never changes settings or sessions, implying it is safe for inspection. It contrasts with the sibling by focusing on analysis and quick wins, but does not provide explicit 'when to use vs when not to use' guidance. The context is clear enough for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v0.1.0- First observed
token_guardian_recommend_route - First observed
token_guardian_usage_snapshot
TDQS
Each tool serves a distinct, non-overlapping purpose: one provides a usage snapshot with improvement suggestions, the other recommends a routing choice for a task. There is no ambiguity between them.
Both tool names follow a consistent 'token_guardian_<verb>_<noun>' pattern using snake_case, with descriptive verbs ('usage_snapshot', 'recommend_route') that clearly indicate their function.
With only 2 tools, the server feels minimal for a domain that might benefit from additional diagnostics (e.g., cost breakdown, model listing) or configuration advice. The count is on the low end of acceptable for a focused utility.
The server explicitly avoids any action tools, being read-only. It covers snapshot analysis and routing recommendations but lacks tools for detailed queries (e.g., by time period, by model) or applying any changes, leaving notable gaps for hands-on usage optimization.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Agent Cost Allocator MCP — multi-tenant LLM cost attribution for chargeback billing. Companion to
Agent Token Budget MCP — hard per-session token + spend cap with signed budget-exhausted
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server that helps AI agents reduce token usage by compressing, summarizing, and managing conversation/context data more efficiently.11MIT
- AlicenseAqualityBmaintenanceMCP server that reduces token consumption in AI coding assistants by up to 90% via structural reads, PreToolUse hooks, and tp-\* subagents.256875MIT
- AlicenseNot gradedqualityDmaintenanceMCP server that analyzes AI agent session logs to find token waste and optimization opportunities.20MIT
- FlicenseNot gradedqualityDmaintenanceAn MCP server that reduces token usage by lazily loading skills and tools only when needed, and routing repetitive subtasks to ML backends instead of the LLM.-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/KitsuneTech1/token-guardian-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server