intent-diff-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@intent-diff-mcpstart a task to add a postgres connection pool"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Intent-Diff MCP
Catch context drift ā the scope an AI coding agent quietly adds beyond what you actually asked for.
An MCP server with two tools. You lock in your original request; the agent, before it declares "done", diffs the actual code changes against that request and reports anything it added on its own ā an unrequested library, a metrics layer, a new abstraction ā so you can Keep it or Roll it back.
The diff is deterministic ground truth (git doesn't lie). The judgment is a fresh Claude call that didn't write the code, so it has no stake in rationalizing the additions.
š¤ Intent Diff Report
Requested: Add a Postgres connection pool in db.ts ā ā
done
[+] Added on my own (drift candidates):
- New metrics.ts with Prometheus instrumentation ā you never asked for
observability or a new dependency (prom-client).
[?] Worth confirming:
- Should metrics.ts stay, or be removed?
ā Keep or Rollback?Tools
Tool | When | What it does |
| A non-trivial feature/refactor is requested | Records your request + snapshots the current git HEAD. Cheap (one SHA). |
| Before saying "done" on substantial changes | Diffs everything since the snapshot, runs the judge, returns a drift report. |
Related MCP server: docverity
Install
Requires Node ā„ 18 and either a logged-in Claude Code session or an ANTHROPIC_API_KEY (see Auth).
Add to your MCP client config (Claude Code .mcp.json, mcporter.json, etc.):
{
"mcpServers": {
"intent-diff": {
"command": "npx",
"args": ["-y", "@ssh00n/intent-diff-mcp"]
}
}
}That's the whole setup ā npx fetches and runs it.
Recommended workflow (hybrid)
Intent-Diff is not meant to gate every edit ā that just taxes trivial work. The sweet spot: always snapshot (cheap), judge on demand (costly). Drop this into your project's CLAUDE.md (a ready copy is in examples/CLAUDE.md):
# Intent-Diff workflow (hybrid)
- When starting a non-trivial feature or refactor, call `start_task` with the
user's request verbatim. (It's cheap ā just a git snapshot.)
- Before declaring a substantial change "done" ā new feature, or edits spanning
multiple files ā run `get_intent_diff` and show the report.
- Treat the report as advisory: surface drift to the user and ask Keep/Rollback.
Skip it for one-line fixes and pure exploration.Prefer softer or stricter? See Tuning the rules.
Auth
The judge needs to call Claude. Two ways, checked in this order:
ANTHROPIC_API_KEY(recommended for CI and contributors) ā normal API billing.Local Claude Code subscription ā if no API key is set, the server reuses the OAuth token your Claude Code login already stores (macOS Keychain or
~/.claude/.credentials.json), refreshing it when near expiry. No extra setup if you're already logged into Claude Code.
Pick the judge model with INTENT_DIFF_MODEL (default claude-sonnet-4-6).
Tuning the rules
The workflow above is the hybrid default. Adjust to taste:
Softer ā drop the
get_intent_diffline; call it only when you ask "check drift". The agent self-checks the rest of the time.Stricter ā make both calls mandatory ("always
start_taskfirst; never say done withoutget_intent_diff") and require an explicit Keep/Rollback answer. Good for team enforcement; higher overhead.
Note: a heavy-handed rule can bias an agent toward under-implementing (skipping necessary error handling for fear of "drift"). The judge is told to ignore refactors and necessary error handling, but keep the rule proportional to the stakes.
How it works
start_task writes <repo>/.mcp/intent_state.json, keyed by git branch, with your
request and the HEAD SHA. get_intent_diff collects git diff <baseSha> plus any
untracked files (skipping binaries, lockfiles, and .mcp/), caps it, and sends
{ original intent, actual diff } to the judge, which returns structured JSON that's
rendered into the report.
Development
npm install
npm run typecheck # tsc --noEmit
npm test # Tier-1 unit tests (no network)
npm run test:smoke # end-to-end via MCP client, stub judge (no network)
npm run eval # judge accuracy over labeled fixtures (needs auth ā real calls)Contributing
Contributions welcome ā see CONTRIBUTING.md. Good first areas: more judge fixtures, additional credential sources, non-git VCS support.
License
MIT Ā© ssh00n
Available Tools
2 toolsget_intent_diffA
Compare the saved original intent against all code changes since start_task, and report context drift (scope the agent added on its own). Call this right before telling the developer the work is done.
| Name | Required | Description | Default |
|---|---|---|---|
| target_dir | No | Optional path inside the target repo. Defaults to the server's CWD. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden. It explains the tool performs a comparison and report, implying no destructive side effects. However, it does not explicitly state that it is read-only or disclose authentication or rate limits, which would improve transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with two sentences: one for purpose and one for usage. No redundant information, every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description adequately implies the tool returns a 'report' of context drift. It covers purpose, timing, and input parameter. Minor gap: no details on output format or structure, but sufficient for an AI agent to understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter 'target_dir' is fully described in the input schema (100% coverage). The description adds no further meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: comparing saved original intent against code changes to detect context drift. It uses specific verbs like 'compare' and 'report', and distinguishes from the sibling tool 'start_task' by implying this tool is used after task start.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs when to call the tool: 'right before telling the developer the work is done.' This provides clear usage guidance with no ambiguity, and the sibling tool context reinforces that this follows start_task.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_taskA
Lock in the developer's original intent before coding starts. Captures the current git HEAD as a snapshot so later drift is measured from here. Call this the moment a new feature/refactor is requested.
| Name | Required | Description | Default |
|---|---|---|---|
| task_description | Yes | The developer's original request, verbatim and complete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must convey behavioral traits. It explains the snapshot mechanism and drift measurement, but omits details on side effects, success/failure responses, or what the tool actually returns. This is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with three sentences, each providing essential information. Purpose is stated upfront, and there is no extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (single parameter, no annotations, no output schema), the description covers core purpose and usage. However, it lacks information about return values and error handling, leaving some contextual gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for task_description ('verbatim and complete'). The tool description adds no additional semantic value beyond what the schema provides, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Lock in the developer's original intent before coding starts.' It specifies a concrete action (capturing git HEAD snapshot) and distinguishes from sibling tool get_intent_diff, which measures drift after starting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance: 'Call this the moment a new feature/refactor is requested.' It implies the context for use and indirectly contrasts with the sibling tool, but does not explicitly state when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v0.1.0- First observed
get_intent_diff - First observed
start_task
TDQS
The two tools serve distinct purposes: start_task captures the initial state, while get_intent_diff compares changes against that state. No overlap or ambiguity.
Both tools follow a consistent verb_noun snake_case pattern (start_task, get_intent_diff), making their actions clear and predictable.
With only 2 tools, the server is thin, but it is focused on a narrow workflow (snapshot and diff). For its limited scope, the count is borderline but acceptable.
The tool set covers the full lifecycle needed: capturing the original intent (start_task) and reporting drift (get_intent_diff). No obvious gaps for its stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for building and testing AI agents with multi-model experimentation and insights.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Driflyte MCP server which lets AI assistants query topic-specific knowledge from web and GitHub.
Monitor MCP servers, API contracts and AI outputs for schema drift. Alerts on breaking changes.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server that enhances AI agents' coding capabilities by providing zero hallucinations, improved code quality, security-first approach, high test coverage, and efficient context management.15271MIT
- AlicenseNot gradedqualityBmaintenanceMCP server that enables coding agents to check documentation claims against source code, detecting drift and suggesting fixes.142MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that provides on-demand safety for AI coding workflows, enabling inspection, review, checkpointing, and rollback of risky actions.211MIT
- AlicenseNot gradedqualityBmaintenanceA self-hosted MCP server that enables AI coding agents to read, edit, search, and run code in local projects with human review loops and policy controls.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ssh00n/intent-diff-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server