Skip to main content
Glama
wfengq

artifactdiff-mcp

by wfengq

ArtifactDiff

A contract change verification gate for humans and AI agents.

Prove that an AI agent changed only the Office/PDF contract terms you authorized — semantically, visually, and cryptographically — before the document is delivered.

ArtifactDiff is local-first and offline. It turns a requested edit such as “change the payment window from 30 days to 45 days” into a frozen policy, verifies the candidate, and writes an immutable Review Bundle that both humans and agents can inspect.

Project status: 0.1.0 alpha. The contract-safe core and signed Golden Path are implemented and covered by more than 900 automated tests. Packaging and public release automation are still in progress.

What it catches

Candidate

Verdict

Delivery gate

Only the exact authorized clause occurrence changed

PASS

May proceed

The edit is semantically allowed but has an unexplained layout change

REVIEW

Blocked until a human approves that finding

A party, amount, date, signature, seal, attachment, or other protected content changed

FAIL

Blocked and not approvable

ArtifactDiff is deliberately fail-closed. Review remains blocking by default and a FAIL finding cannot be approved away.

Related MCP server: Proof Layer MCP

Three-minute Golden Path

Prerequisites: Python 3.11+ and a local checkout of this repository. The demo uses only synthetic contracts and makes no network request after dependencies are installed.

python -m venv .venv
# Activate .venv using your shell, then:
python -m pip install -e ".[dev]"
python scripts/run_contract_golden_path.py --output build/contract-golden-path
python -m json.tool build/contract-golden-path/summary.json

The run creates three independently verified, Ed25519-signed Review Bundles:

authorized    raw=pass    effective=pass
review        raw=review  effective=review
unauthorized  raw=fail    effective=fail

The review case stays blocked. To demonstrate the signed event-chain transition with an explicitly simulated synthetic human reviewer, use a different output directory:

python scripts/run_contract_golden_path.py --output build/contract-golden-path-approved --approve-review

--approve-review is demo-only. Real approvals remain interactive human actions through artifactdiff review or artifactdiff approve; agents cannot approve their own findings. See the Golden Path walkthrough for the generated files and privacy boundaries.

The workflow

Human instruction
      │
      ▼
contract-safe policy ── freeze + optional Ed25519 authorization
      │
      ▼
controlled agent edit session
      │
      ▼
semantic rules + protected entities + visual envelope
      │
      ├── PASS ───────────────────────────────► deliver
      ├── REVIEW ─► signed per-finding review ─► deliver or reject
      └── FAIL ───────────────────────────────► reject
                          │
                          ▼
                 immutable Review Bundle

The Review Bundle is the system of record. The local review desk is only an authenticated loopback interface over that bundle.

Human and automation interfaces

The same application boundary powers every interface:

  • CLI: create/validate/seal policies, open controlled sessions, verify changes, inspect findings, approve findings, and verify or pack bundles.

  • Local review desk: inspect a bundle and sign one finding at a time without sending contracts or private keys to browser JavaScript.

  • MCP: bounded inspect/draft/validate/local-verify tools for AI agents, with separate input and output roots and no approval or verified-signing capability.

  • Offline HTML: portable human review reports linked to verified facts.

artifactdiff --help
artifactdiff policy --help
artifactdiff verify --help
artifactdiff review path/to/review-bundle

For an MCP host, configure absolute, path-separator-delimited roots before starting the stdio server:

ARTIFACTDIFF_MCP_INPUT_ROOTS=/absolute/contracts
ARTIFACTDIFF_MCP_OUTPUT_ROOTS=/absolute/artifactdiff-output
artifactdiff-mcp

MCP intentionally cannot approve findings or perform verified signing.

Why this is not another PDF diff

Ordinary document diff

ArtifactDiff

Shows everything that changed

Proves whether changes match a pre-authorized instruction

Text or pixels are the final result

Semantic, protected-entity, occurrence, metadata, and visual rules combine into one gate

A screenshot/report is the evidence

Content-addressed Review Bundle with policy, facts, verdict, evidence, signatures, and append-only decisions

Designed only for a person looking at two files

CLI, Python, MCP, offline report, and human review desk share one truth model

“Looks fine” may pass

REVIEW and unavailable evidence block by default

Security and privacy model

  • Local-first and offline; no account or cloud service is required.

  • contract-safe is enabled by default, with no implicit authorization.

  • Exact expected edits are bound to their clause and occurrence.

  • Parties, money, currencies, dates, durations, percentages, headers, footers, signatures, seals, and attachments are protected by default.

  • Offline Ed25519 is the core trust mechanism; enterprise identity can be added through adapters later.

  • Evidence defaults to minimal; full is explicit and sealed is encrypted for named recipients.

  • MCP responses are bounded and do not return full contracts, page images, or key material.

Read the approved verification-gate design and the Plan 4.5 acceptance record for the full trust and verdict model.

Current format support

  • PDF → PDF

  • DOCX → DOCX

  • DOCX → PDF

  • English, Chinese, and bilingual contract structure

  • Optional DOCX rendering through LibreOffice; PDF rendering is local

Development

python -m pip install -e ".[dev]"
python -m pytest -q

The repository currently contains unit, integration, tamper, CLI, MCP, review-desk, cross-format, and signed Golden Path acceptance coverage.

Next milestones

  • Reproducible public synthetic corpus and zero-false-pass gate

  • GitHub Action for binary-document pull request review

  • Demo media and downloadable example Review Bundle

  • Schemas, threat model, SBOM, provenance, and signed alpha release

  • Plugin boundaries for enterprise parsers, renderers, signing, encryption, and storage

ArtifactDiff is being built around one narrow promise: when an agent edits a contract, you can prove exactly what it was allowed to change — and that it changed nothing else.

Available Tools

10 tools
compare_documentsD
ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo
visualNo
after_pathYes
output_dirNo
before_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_contract_policyD
ParametersJSON Schema
NameRequiredDescriptionDefault
afterYes
anchorYes
beforeYes
headingYes
rule_idYes
clause_labelYes
ancestor_pathNo
baseline_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_review_findingD
ParametersJSON Schema
NameRequiredDescriptionDefault
finding_idYes
bundle_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_contractD
ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_documentD
ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
forceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_review_findingsD
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNo
bundle_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seal_local_policyD
ParametersJSON Schema
NameRequiredDescriptionDefault
output_pathYes
policy_pathYes
baseline_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_contract_policyD
ParametersJSON Schema
NameRequiredDescriptionDefault
policy_pathYes
baseline_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_contract_changeD
ParametersJSON Schema
NameRequiredDescriptionDefault
visualNo
output_pathYes
baseline_pathYes
candidate_pathYes
sealed_policy_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_review_bundleD
ParametersJSON Schema
NameRequiredDescriptionDefault
bundle_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 10 tool updatesv0.1.0
    • First observedcompare_documents
    • First observeddraft_contract_policy
    • First observedget_review_finding
    • First observedinspect_contract
    • First observedinspect_document
    • First observedlist_review_findings
    • First observedseal_local_policy
    • First observedvalidate_contract_policy
    • First observedverify_contract_change
    • First observedverify_review_bundle

TDQS

D1.9/5.0
Disambiguation2/5

The tool set includes several pairs with ambiguous boundaries: inspect_document vs inspect_contract likely overlap, and compare_documents vs verify_contract_change seem to address similar tasks. Without descriptions, an agent would struggle to distinguish which tool to call for a given operation.

Naming Consistency5/5

All tools use a clear snake_case verb_noun pattern (compare, inspect, draft, validate, seal, verify, list, get). Consistency is high and follows a predictable convention.

Tool Count5/5

10 tools is within the ideal range for a domain covering document comparison and contract policy management. The count is well-scoped with no obvious bloat.

Completeness4/5

Core workflows for contract policy drafting/validation/sealing and review finding retrieval are covered. However, there is no tool for listing or managing documents/contracts, and the generic artifact diffing implied by the server name is thin, leaving minor but noticeable gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to analyze Ethereum wallets, simulate transactions, and draft transfers with deterministic policy and risk scoring, requiring human approval before on-chain execution.
    2
    ISC
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides cryptographic governance receipts for AI agents, enabling pre-execution evaluation and signed verdicts (EXECUTE/BLOCK/REVIEW/SHADOW) with offline-verifiable audit trails.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to create cryptographically verifiable receipts of their delegated work, with capabilities for multi-party approval and offline verification.
    11
    162
    Apache 2.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/wfengq/artifactdiff'

If you have feedback or need assistance with the MCP directory API, please join our Discord server