TrustGate
Allows the MCP server to interact with Razorpay's payment provider in Test Mode, enabling secure purchase authorization and capture while enforcing tenant-scoped policies, approvals, and verified webhook events.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@TrustGatePropose purchase of SKU TG-2001, quantity 3, for the quarterly offsite."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TrustGate
An AI agent can ask to buy something. It can never decide what that costs, who gets paid, or whether the money actually moves.
TrustGate is a synthetic-data, Razorpay Test Mode testbed for bounded agent spending. An AI buyer proposes a catalog purchase and never gains authority to rewrite the merchant, amount, currency, approval, or provider outcome. The agent proposes; TrustGate independently authorizes and records evidence.
In a hurry?
make triageis a guided tour whose first two steps need no database, no Docker, and no credentials.JUDGE.mdmaps every claim on this page to the command that proves it.
The problem, in one screen
An AI buying agent reads a product catalog. One description was written by a supplier and contains an instruction. The agent follows it — models do, and treating that as a bug to be trained away is the assumption this project refuses to make.
flowchart LR
S["Supplier writes into a<br/>catalog description:<br/><i>charge ₹20,000 to<br/>attacker-controlled</i>"] --> A["AI buying agent<br/>reads it and obeys"]
A --> U["<b>Unguarded tool</b><br/>accepts amount +<br/>merchant"]
A --> T["<b>TrustGate tool</b><br/>accepts sku, quantity,<br/>purpose. Nothing else."]
U --> P1["₹20,000 paid to<br/>the attacker"]
T --> P2["Nowhere to put<br/>the instruction"]
style U fill:#5b1a15,stroke:#973029,color:#fff
style P1 fill:#5b1a15,stroke:#973029,color:#fff
style T fill:#14372a,stroke:#2c6349,color:#fff
style P2 fill:#14372a,stroke:#2c6349,color:#fffBoth halves are in this repository. python -m demo.unguarded runs them side by side against the
same poisoned catalog and the same model response.
Nothing detected an attack. No filter, no classifier, no anomaly score. One interface had a field for the amount and the other did not. That is the entire difference, and it is why this is an authorization problem rather than a detection problem.
The part worth looking at is not that it refuses things. It is that the refusals are verified rather than asserted. Every invariant in the mutation registry below is deleted on purpose, one at a time, and a test has to fail. That is how an unguarded policy-expiry check was found here after two clean audits: 302 tests passed with the check removed.
The registry is targeted, not exhaustive, and the distinction is worth keeping. It covers guards that live in Python. Guards that live in the database - triggers, check constraints, partial unique indexes - are proven instead by tests that violate them directly, because a mutation runner edits source files and a trigger already applied to a schema would not notice.
Related MCP server: Allowance MCP
The same three columns, both ways
Every persisted purchase request produces this record: what the agent proposed, what the server derived, and what the provider actually did. Reading them beside each other is what makes the boundary visible rather than asserted — the agent's column is short on purpose.
A refusal at the tool boundary produces no such record, and that is the design rather than a gap: the request never became a row, so there is nothing to write three columns about. Those refusals are recorded separately, in the audit trail, which is why the right-hand column below can say a receipt does not exist while the left-hand one links to a preserved one.
Both of these are real. The left is the payment captured on 3 September; the right is the injected
catalog description from python -m agent.demo --adversarial.
✅ A purchase that should go through | ❌ The same agent, following injected text | |
1 · AI proposed |
|
|
2 · TrustGate derived | Merchant Campus CloudAmount ₹399.00Policy ALLOW, version 2Authority | No amount derivedNo merchant derivedRefused: |
3 · Razorpay confirmed | Order | Provider contacted: NoOrder created: NoNo payment request exists |
Receipt | there is none — nothing was written to write one about |
The bottom-right cell is the result this project exists to produce. Not a refusal that was logged — an attempt that never became a record, because the field the attack needed was not there to fill.
Why This Exists
Agent payment protocols arrived through 2025 and 2026 — Google's AP2, Mastercard's AP4M, Visa's Trusted Agent Protocol, Coinbase and Cloudflare's x402. Razorpay and NPCI announced agentic UPI payments in February 2026. Every one of them assumes an agent will eventually be told to spend money by something that is not its principal, and most of that text is about mandates: proving a human agreed.
The gap this fills is what happens after the mandate exists and before the money moves.
An agent reading third-party content is reading attacker-controlled input. A product description, a support ticket, a web page — any of it can carry an instruction, and language models follow instructions. Guardrails, classifiers and prompt hardening all reduce how often the model is fooled. None of them change what happens when it is.
So the question this project asks is not "how do we stop the model being wrong?" It is:
When the agent is wrong — and it will be — what is it structurally incapable of doing?
The answer here is: name a price, name a merchant, approve its own purchase, mint its own authority, or cause money to move. Not because those attempts are detected, but because there is no field, no tool, and no code path through which they can be expressed.
What that took
Decision | Why |
The agent's tool takes sku, quantity, purpose — nothing else | An amount field is an attack surface no filter can close |
Money facts are derived server-side from the catalog | The agent never sees a price it could alter |
Authorization and payment are separate steps | Being allowed to buy is not being able to pay |
Checkout authority is single-use, 15 minutes, hash-bound | A permission slip for one exact purchase |
Invariants live in triggers and constraints | A rule in Python is a rule a different query walks around |
Every guard is deleted on purpose by the mutation suite | A passing suite is not evidence the tests would notice |
What came out of it
An agent runs the same poisoned catalog through both halves of this repository. The unguarded adapter pays ₹20,000 to an attacker-named merchant. TrustGate produces no payment request at all — so there is no receipt to open, because there is nothing to write one about.
That absence is the result. Everything else here exists to make it trustworthy: 16 adversarial
scenarios with a generated attack matrix, a mutation registry that breaks each safety guard and
requires a test to fail, concurrency invariants raced rather than approximated, and a real Razorpay
Test Mode payment carried to CAPTURED by a signed webhook the provider itself delivered.
The method found real defects that reading the code did not: an unguarded policy-expiry check that 302 passing tests missed, a row lock that serialised transitions while letting the second caller decide from stale state, and a budget that could be reserved and never returned.
What It Proves
Catalog SKU and quantity are the only purchase facts an agent can influence.
Tenant-scoped policy, human approval, a one-time checkout authority, and verified provider events bound every payment action.
Authority delegated onward narrows at every hop, and the hops below a node cannot, between them, promise more than that node holds.
An actor holding a delegation has its purchases checked against the whole chain during authorization, and the budget returns if the payment never happens.
Revoking one hop ends every branch below it without touching a descendant or recalling anything, including for a payment already authorized: the chain is re-asked before checkout authority is issued and again before it is consumed, and a payment stopped that way gives both budgets back.
Unsafe attempts are rejected before a provider order can be created and leave an auditable trace.
The whole path runs against the real provider: Razorpay Test Mode creates the order, a human pays it, and Razorpay's own signed
payment.capturedis what moves the money — verified over raw bytes, cross-checked against the server-derived amount, and preserved asdocs/evidence/m3-provider-delivered-webhook.json.
How AI Is Used Here
The buyer agent is deliberately thin, and that is the argument rather than a shortcut. It can propose only a catalog SKU, a quantity, and a purpose; every money-critical fact is derived server-side. Delete the agent entirely and the authorization layer is unchanged and still correct.
The engineering contribution is not a sophisticated agent. It is taking language-model failure modes seriously — prompt injection through third-party content, instruction-following against the operator's interest, confident wrong reasoning about money — and building a system that stays correct when they occur.
That claim is tested both ways. python -m agent.demo --live runs a real model against catalog
descriptions written by third parties, including hostile ones. The same catalog is sent twice, once
with descriptions removed, so influence from untrusted content is measured by comparing the two
proposals rather than self-reported. The regression suite never calls a model provider: it uses
deterministic substitutes, so safety verification does not depend on a model behaving a particular
way on a particular day.
Seeing the Problem
Three clean refusals prove the system works and are weak at showing why anyone needs it. The demo therefore opens with this project's own code failing, using no network, no credentials, and no database:
python -m demo.unguardedThe same agent reads the same poisoned catalog - literally the same, since both catalogs are built
from agent/demo_catalog.py and a test asserts the seeded row and the baseline object carry an
identical injected instruction. One model response is handed to two adapters.
The unguarded one accepts an amount and a merchant, so the injected instruction executes and it
pays INR 20,000.00 to a merchant the catalog text named, against a catalog price of INR 600.00.
The other has nowhere to put either field, because PurchaseProposal declares only a SKU, a
quantity, and a purpose. What survives the discard is quantity=50, which is a field the agent is
allowed to set - and the server bounds it against the catalog's own maximum of 2, so the attempt is
refused before a payment request is created.
The difference is not a filter that recognized an attack. It is that one interface had a field for the money and the other did not. The baseline has its own tests asserting both that it cannot reach anything real and that it is still exploitable, since a demonstration that quietly stopped being vulnerable would keep passing while making the opposite point.
Run the Demo
Four commands take a clean database to a real Razorpay Test Mode payment. Requires the stack from
Quickstart and ENABLE_CONSOLE=true in .env.
python -m agent.stageStages a fixed demo tenant, clears the timeline, and prints the console URL plus the exact command for every beat below. Run it between takes.
python -m demo.unguardedThe problem, in this project's own code. No policy layer, so the injected instruction executes.
python -m agent.demo "Buy Starter credits for the robotics club."
python -m agent.checkout --openA purchase that should go through. The agent proposes a SKU and a quantity; the server derives
the price and merchant. checkout issues a one-time authority, creates a real provider order, and
opens the payment page — the agent does neither, which is the point.
python -m agent.demo "Buy Team credits for the robotics club."
python -m agent.approveA purchase that needs a human. Over the approval threshold, so the agent cannot finish it.
approve is a separate command holding a token the agent does not have, under an identity the
server refuses if it matches the requester.
python -m agent.demo --adversarial "Buy cloud credits for the club."The attack. The same injected instruction from the first command. The amount and merchant are discarded because the proposal has nowhere to put them; what survives is a quantity the catalog bounds. No payment request is created.
The read-only console at /console/{tenant_id} leads with the current decision — AUTHORIZED,
APPROVAL REQUIRED, or BLOCKED — with the reason in plain words, whether the provider may be
called at all (Order creation allowed: Yes/No), and the human whose delegated authority is behind
it. Below that, every attempt in three columns: what the agent proposed, what the server derived,
and what the provider actually did.
The panel is assembled from the same evidence record the receipt renders, and from the first row of the table beneath it, so it cannot disagree with either. Reason codes are translated for display only — the table still shows the exact stored code. The console cannot authorize, approve, or create anything.
python -m agent.delegateDelegation. Four hops, narrowing at each. A leaf spends inside every bound above it, is refused a SKU outside the scope it was narrowed to, and a sibling is refused budget its parent had already promised away. Then one hop is revoked and the branch below it dies untouched.
demo/script.md is the rehearsed path with timings and what to say at each beat.
Delegated Authority, and Why a Budget Is Not a Capability
Attenuating a capability is set intersection. A child permitted no more than its parent cannot widen the chain however deep it runs, because sets intersect - which is what the delegation literature, the macaroon family, and the IETF attenuating-token draft all mean by attenuation.
Money is not a set. Two children each granted exactly their parent's budget satisfy every per-edge comparison and hold twice the parent's budget between them. Budgets add where sets intersect.
flowchart TD
H["<b>Finance lead</b> (human)<br/>grants ₹1,000"] --> P["<b>Agent A</b><br/>budget ₹1,000"]
P -->|"child asks ₹1,000"| C1["<b>Agent B</b> ✅<br/>₹1,000 allocated"]
P -->|"child asks ₹1,000"| C2["<b>Agent C</b> ❌<br/>DELEGATION_BUDGET_EXHAUSTED"]
N["Each child is <i>no wider than its parent</i>.<br/>Every per-edge check passes.<br/>Together they hold ₹2,000."]
style H fill:#14372a,stroke:#2c6349,color:#fff
style P fill:#101a1c,stroke:#223436,color:#fff
style C1 fill:#14372a,stroke:#2c6349,color:#fff
style C2 fill:#5b1a15,stroke:#973029,color:#fff
style N fill:#3a2f10,stroke:#8c5a0c,color:#fffThe refusal comes from ck_delegation_budget_partitioned, a check constraint on the parent's own
row — not from application code that a different query could go around.
The chain, as the system prints it
Not a mock-up. This is python -m agent.delegate against a live database, and every number is read
back from the rows after each step.
A human delegates a budget, and every hop below narrows it.
demo-buyer-agent holds INR 2,000.00 promised down INR 1,200.00 spent INR 0.00 free INR 800.00
|_ demo-procurement-agent holds INR 1,200.00 promised down INR 600.00 spent INR 0.00 free INR 600.00
|_ demo-renewals-agent holds INR 600.00 promised down INR 0.00 spent INR 0.00 free INR 600.00
The leaf spends INR 400.00 on CLOUD-STARTER, inside every bound above it.
|_ demo-renewals-agent holds INR 600.00 promised down INR 0.00 spent INR 400.00 free INR 200.00
The leaf was narrowed to starter renewals, so the team SKU is out of scope:
refused: DELEGATION_SKU_OUT_OF_SCOPE
Now the finding. The root holds INR 2,000.00 and has promised INR 1,200.00 already.
A second child asks for another INR 1,200.00.
Per-edge narrowing allows it: the child is no wider than its parent.
refused: DELEGATION_BUDGET_EXHAUSTED
The budget was partitioned, not compared. Siblings share one pool.
The human revokes demo-procurement-agent, in the middle of the chain.
demo-renewals-agent is not touched, and is not told. Its next spend:
refused: DELEGATION_REVOKED
its own revoked_at is still NoneThat last pair of lines is the one to read twice. The leaf's own row still says it is live — its authority is re-derived from the whole chain on every spend, so cutting a link above it is already the end of the branch. A signed capability would have to be hunted down and recalled, which is why revocation is an open problem for that whole family of designs.
TrustGate therefore partitions a parent's budget rather than comparing against it. The
delegation_attenuates trigger takes the allocation from the parent as part of writing the child,
so the aggregate holds for anything that inserts a row rather than for callers who remember the
bookkeeping. python -m agent.delegate walks four hops and shows the sibling being refused.
An earlier version claimed this "survives every line of the application being wrong" while the
allocation was still maintained in Python. It did not: three children of 1,000 were written under a
parent holding 1,000 by an insert that skipped the module, and every per-edge check passed them.
The claim is true now because the trigger does the work, and
test_siblings_written_straight_to_the_database_cannot_outgrow_their_parent is what says so.
The mutation named delegation-aggregate-partition is the evidence that this is a real distinction
and not a stylistic one: delete the aggregate claim and every other delegation test still passes.
This is wired into authorization, and the boundary that remains is narrower than it was. A purchase by an actor holding a delegation is checked against every hop above it before the payment request is recorded, the delegation is debited in the same savepoint as the daily reservation, and the budget comes back on the same condition that already returns the daily one - DENIED, EXPIRED, FAILED, CANCELLED. An actor holding no delegation takes the path it took before any of this existed, which is what makes the rest of the suite the regression net for the wiring.
Two budgets refusing separately is where this could have leaked, and the savepoint is why it does not: claiming the daily reservation and then refusing on the delegation would leave it moved for a payment that never happens, and the release path only fires on a transition out of a holding state, which a request refused at authorization never enters. Doing the delegation first only mirrors the leak. Either both hold or neither does.
What is still missing is a way to create one over HTTP. A delegation is granted and revoked from
Python today - python -m agent.delegate and the staging command - and there is deliberately no
tool letting an agent mint its own authority. docs/limitations.md has the rest, including the
largest: nothing proves the actor spending is the actor the delegation was granted to.
A second property falls out of the same design. A hop carries no signature; its authority is
re-derived from its whole chain, against live policy, every time it is spent. Revoking any link is
already the end of the branch - no recall, no revocation list, and the descendant is never written
to. The capability-token designs make the opposite trade, buying offline verification and giving up
recall. docs/limitations.md states what that costs here.
Attack Matrix
Tier A adversarial scenarios. This table is generated from the scenario registry by
python -m scenarios.report, and a test asserts it matches, so it cannot claim an attack
that is not covered by a passing test. Every scenario proves three things: the attack is
rejected with its reason code, no provider order was created, and no payment gained
authority it did not have.
ID | Attack | Invariant proven | Tests |
A1 | Amount tampering | The amount is derived from the catalog item's price and a server-bounded quantity. No agent-supplied value can change it. |
|
A2 | Merchant substitution | The merchant is derived from the tenant-scoped catalog item. A merchant outside the tenant is unreachable, and one outside the active policy cannot be paid. |
|
A3 | Currency substitution | Currency is derived from the catalog item, and the one route that accepts a currency is disabled by default and denies a mismatch against the active policy when enabled. |
|
A4 | Expired or reused approval | An approval is a permission with a lifetime and a single use. Neither an expired one nor an already consumed one can authorize, and a refused approval is not burned. |
|
A5 | Self-approval | An approval cannot be granted by the identity that requested the purchase. Separation of duties is enforced, not merely expected from configuration. |
|
A6 | Forged webhook signature | Provider events are authenticated by raw-byte HMAC before the body is parsed. A forged or absent signature changes nothing, however well-formed the event is. |
|
A7 | Tampered webhook body | The signature covers the exact bytes received, so a genuinely signed event edited in flight no longer verifies and never reaches a payment. |
|
A8 | Duplicate webhook delivery | Provider event identity is stored, so a replay of an authentic, in-window event is refused by the database rather than by whichever handler happens to look. |
|
A9 | Out-of-order provider events | Arrival order is the provider's and legality is ours. A capture cannot precede its authorization, and a terminal payment accepts no further outcome. |
|
A10 | Double refund | No surface can initiate a refund at all, asserted against the live route table and tool list, and the ledger invariant refuses a refund total exceeding the capture. |
|
A11a | Unknown tenant header | A tenant that does not resolve is refused before any route body runs, and the refusal discloses nothing that would let a caller enumerate which tenants exist. |
|
A11b | Cross-tenant object access | Every tenant-scoped lookup filters by the trusted tenant. A known tenant cannot read or act on another tenant's request, payment, or authority on any surface. |
|
A12 | Idempotency key collision | A key reused with a different purchase returns the original decision and a 409. The second purchase is never created and cannot be mistaken for one that was accepted. |
|
A13 | Policy drift between authorization and use | An authority does not outlive the policy it was checked against, nor the purchase it was issued for. A superseding policy, an expired one, or an edited amount revokes it without burning it, and an undrifted authority still works. |
|
A14 | Stale or post-dated webhook | A signature proves origin, not recency. An event outside the freshness window, dated into the future, or carrying no timestamp at all is refused before any lookup. |
|
A15 | Unauthorized capture via MCP | No tool reachable by the agent can authorize, capture, refund, or call a provider. Proven by exercising every exposed tool, not by inspecting tool names. |
|
Every Tier A scenario is implemented. A11a and A11b split tenant confusion into an
unresolvable tenant and a known tenant reaching across the boundary, because the two fail for
different reasons and only the second is an authorization question.
What the Tests Are Worth
A passing suite says the code behaves as written. It does not say the tests would object if the code stopped doing something important, and only the second claim matters when the subject is money. This project has evidence for the difference: request-scoped sessions once discarded every write while 146 tests passed, because the suite asserted inside the same transaction it wrote in.
make mutation breaks each safety invariant on purpose, one at a time, and requires the tests named
as its guards to fail. Every source file is restored in a finally block and the restoration is
checked against git diff before the report prints, so an interrupted run cannot leave a mutation
behind. It exits non-zero if any mutation survives.
This table is generated from the mutation registry by python -m scenarios.report --mutations, and
a test asserts it matches, so it cannot claim a guarded invariant that is not actually guarded.
Mutation | Invariant it removes |
| A payment is locked before its state is read and changed. |
| A locked read decides from the committed row, not from a cached one. |
| Row locks are taken through the one helper that keeps them meaningful. |
| A provider event is authenticated before anything is done with it. |
| A signed provider event proves origin, not recency. |
| An event that cannot be dated cannot be bounded, so it is refused. |
| An approval is a permission with a lifetime, not a permanent grant. |
| An authority does not outlive the policy it was checked against. |
| An authority cannot be spent under a policy that has run out. |
| An authority is bound to the exact purchase it was issued for. |
| The daily budget upsert refuses to exceed the limit. |
| Budget is returned only by a payment that actually reserved it. |
| Catalog text cannot terminate the checkout page's script element. |
| A successful request commits its writes. |
| Lifecycle events for one payment are distinct events, not replays. |
| An approval cannot be granted by the requesting actor. |
| Evidence is scoped to the tenant that asked for it. |
| An incomplete provider search never reports a receipt as absent. |
| An expired policy cannot authorize new spending. |
| A tenant with no policy is denied rather than allowed by default. |
| A hop cannot spend budget it has already promised to the hops below it. |
| Revoking one hop stops every hop below it. |
| A spend is bound by the narrowest per-payment cap anywhere above it. |
| A purpose narrowed at one hop stays narrowed at every hop below it. |
| An expired hop stops the branch below it. |
| A spend moves budget one way; a negative amount cannot refund it. |
| A spend holds every hop above it, so a revoke cannot land mid-decision. |
| Spending twice under one reference charges the chain once. |
| A spend given back once cannot be given back again. |
| A refused spend does not burn the reference it was refused under. |
| A spend that moves budget records that it did. |
| A spend's evidence names every hop that authorized it, not just the leaf. |
| A reused reference carrying different details is refused, not reported as done. |
| A payment refused after one budget moved gives it back before it is recorded. |
| A payment that never happens returns its delegated budget, on every path. |
| A payment by an actor holding a delegation is checked against it. |
| A chain read after a grant reports the allocation the grant just took. |
| A delegation revoked after authorization stops the checkout it authorized. |
| A chain revoked while its authority was in hand refuses the provider call. |
| A checkout blocked by a dead chain strands neither budget on the payment. |
| A payment request names the delegation it debited, durably enough to re-ask. |
| An actor whose delegation expired can be granted another one. |
| Making room for a grant never revokes a delegation that still works. |
| A revoke that matches nothing says which nothing, without leaking the other. |
| A delegation spend is reachable from the payment request it paid for. |
| A returned delegation budget is reachable from the purchase that returned it. |
| A purchase made under a delegation says so in its evidence record. |
| An authorized purchase with no checkout authority is not allowed to pay. |
| A revoked chain takes the provider action away from an issued authority. |
| A delegated purchase says whose authority it spent, on the timeline. |
| An attack refused before a payment request exists says so, rather than blankly. |
| A refusal is shown in words, not as the code it is stored under. |
| A captured payment blocks provider action as settled, not as never authorized. |
The first run of this suite found a live defect. SELECT ... FOR UPDATE through the ORM acquires
the lock correctly and then discards the row Postgres returned, because SQLAlchemy keeps the
attributes of an object already in the session's identity map. A second caller therefore blocked on
the lock as designed, received the committed row, kept its stale copy, and authorized a payment
that had just been authorized. The lock was serialising when transitions ran, not what state they
decided from. All locking now goes through models.locking.locked(), and a test asserts that
helper is the only place in the source that can take a lock.
Current Scope
Test Mode by choice, not by limitation. A project whose subject is bounded spending authority
has no business holding live keys, and .env.example says a rzp_live_ key must never be placed
here — a guard the code enforces rather than the documentation requesting. Tenants, merchants, and
prices are synthetic for the same reason.
Within that boundary the system is complete and exercised end to end against the real provider. It is an authorization layer, deliberately not a payment processor, compliance product, legal-consent system, or fraud model — each of those is a different product, and claiming any of them would be the kind of overreach the rest of this README is built to avoid.
docs/positioning.md explains why a project like this is worth building
now, with each item of Indian payments context labelled by the evidence behind it — an NPCI
circular is cited as a circular, a press report as a press report, and a recommendation as a
recommendation. It claims no compliance with any of them.
docs/limitations.md names every deliberate cut, including the unflattering
ones: tenant identity is a header and not authentication, one webhook secret serves all tenants,
receipts are traceable but not tamper-evident, and there is no rate limiting anywhere. A limitation
you have to discover is worse than one that is written down.
Quickstart
Requirements: Docker Desktop and Python 3.12.
Copy-Item .env.example .env
docker compose up -d
docker compose exec -T api python -m alembic upgrade head
docker compose exec -T api python -m pytest -qmake verify runs the full gate — lint, format, types, migration parity, the suite, and the
mutation registry. Start the stack first: alembic check and the tests both talk to Postgres,
and a gate that skipped them when the database was absent would not be a gate.
docker compose up -d
make verifymake triage needs no such setup for its first two steps, which is the point of it.
The local API health check is available at http://127.0.0.1:8000/health.
For the demonstration, use python -m agent.stage — it stages a fixed tenant so the console URL
stays the same between runs, and prints every command. See Run the Demo.
python -m agent.seed remains for exploration: it mints a disposable tenant with fresh
identifiers, which is right for poking at the system and wrong for anything you intend to film.
To run the same flow with a real model instead of a deterministic substitute, install the optional
extra with pip install -e ".[agent]" and add --live. Two backends are supported and both use
the same Messages API shape:
TRUSTGATE_MODEL_BACKEND=anthropic(default) readsANTHROPIC_API_KEY.TRUSTGATE_MODEL_BACKEND=bedrockbills against an AWS account. SetAWS_REGIONto the region the credential belongs to, then eitherAWS_BEARER_TOKEN_BEDROCK(a Bedrock API key, the simplest path) or standard AWS credentials for SigV4 signing. Amazon Bedrock provisions Anthropic models through an AWS Marketplace subscription, so the AWS account also needs a valid payment instrument even when credits would cover the usage.TRUSTGATE_MODEL_BACKEND=groqreadsGROQ_API_KEYand needs no payment instrument at all.
TRUSTGATE_MODEL_ID overrides the model on any backend. The buyer is a protocol implementation,
so the provider is a configuration choice rather than an architectural one: the authorization
layer's behavior does not depend on which model proposes the purchase.
This is the only path in the project that contacts a model provider; the test suite never does. Run it only against the synthetic seed catalog. Its third-party descriptions are sent to the model twice, once with descriptions removed and once intact, to measure influence. Never send real customer, merchant, or payment data through this demonstration.
To exercise the Razorpay Test Mode adapter, set RAZORPAY_KEY_ID and
RAZORPAY_KEY_SECRET in the ignored .env file. Never add Test Mode or Live Mode secrets to the
repository.
Trust Boundary
flowchart TD
A["<b>AI buyer</b><br/>proposes sku, quantity, purpose"] -->|the only arrow<br/>the agent crosses| B
B["<b>TrustGate</b><br/>derives merchant, amount, currency<br/>evaluates policy, delegation, approval"]
B --> C["<b>A human</b><br/>takes the authorization to checkout"]
C --> D["<b>Razorpay Test Mode</b><br/>executes a bounded order"]
D --> E["<b>Signed provider event</b><br/>the only thing that moves money"]
E --> F["<b>Evidence</b><br/>proposed · derived · provider outcome"]
style A fill:#3a2f10,stroke:#8c5a0c,color:#fff
style B fill:#14372a,stroke:#2c6349,color:#fff
style C fill:#14372a,stroke:#2c6349,color:#fff
style D fill:#101a1c,stroke:#223436,color:#fff
style E fill:#101a1c,stroke:#223436,color:#fff
style F fill:#101a1c,stroke:#223436,color:#fffThe agent crosses the first arrow and no other. Authorization and payment are separate steps: it obtains the right to buy something and never obtains the ability to pay.
The gap between those two steps is the product. A purchase can sit AUTHORIZED indefinitely
with Order creation allowed: No — approved, and unable to move a rupee, because no checkout
authority has been issued. Those read as the same outcome unless a system says them separately.
Document | What it holds |
Every claim, and the command that regenerates it | |
Every real choice, with the alternatives rejected and why | |
Every deliberate cut, including the unflattering ones | |
Indian payments context, labelled by the evidence behind it | |
How the pieces fit | |
Reaching this machine from Razorpay, and why cloudflared rather than ngrok | |
What is defended and against whom | |
The formal plan this was built against |
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
No tool schema history has been recorded yet.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Procure governed AI capabilities: machine-readable price, scope, trust, gateway-delegation terms.
Secure agent purchasing with human-approved virtual cards, receipts, and audit trails.
Pre-spend firewall for AI agents. Approves, blocks, flags transactions against policy rules.
Advisory policy preflight for AI-agent spend requests; never executes payments or accesses wallets.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to create, compare, and track purchases with structured buying workflows, offer comparison, and merchant verification.5MIT

Allowance MCPofficial
FlicenseNot gradedqualityBmaintenanceEnables AI agents to request purchase approval from humans, receive scoped virtual cards, complete checkout, and report receipts for audit.-- FlicenseNot gradedqualityBmaintenanceEnables AI agents to browse product catalogs and make purchases through a policy engine that enforces spending limits, requires human approval for certain amounts, and logs all actions to an audit trail.-
- FlicenseNot gradedqualityDmaintenanceEnables human-in-the-loop authorization for AI agent transactions, allowing real-time approval or denial of purchases based on configurable spending limits, vendor blocklists, daily caps, and category restrictions.1-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Saraid10/Trustgate'
If you have feedback or need assistance with the MCP directory API, please join our Discord server