Skip to main content
Glama
aks129

HealthClawGuardrails

by aks129

HealthClaw Guardrails

The open-source security layer between AI agents and clinical data.

FHIR standardized how health data is structured. MCP standardized how AI connects to tools. Nobody standardized the guardrails in between. This project does.

Release License CI Code size

Stars Forks Issues Contributors Last commit

Tests MCP tools FHIR Guardrail conformance Glama score Python Docker

Quick Start · MCP Tools · Recipes · Roadmap · Claude Plugin · Architecture · healthclaw.io · Contributing · Dev Guide


What it is: an open reference implementation of the FHIR × MCP guardrail layer — PHI redaction, immutable audit, step-up auth, and tenant isolation — that sits between any AI agent and any FHIR server. Built in the open as a community project, MIT-licensed. Not a product, not a pitch: if the pattern is useful, take it; if it's wrong, tell us or fix it.

This is a community effort. It's most useful when implementers, clinicians, and standards folks poke holes in it. Issues, PRs, and "you got the SDC extraction wrong" critiques are all welcome — start with CONTRIBUTING.md and the Code of Conduct.

At a glance: v1.10.0, with 1,490+ Python and 170 Node tests across 29 MCP tools. CareAgents is the hosted consumer app: passkey sign-in, advisors, and web/Telegram/iMessage. Two rails run end to end — real-world actions behind a provably out-of-band gate, and forms ($populate → human review → provenance PDF). Standards: FHIR R4 US Core v9 and R6 v6.0.0-ballot3, HL7 SDC forms, NQF 0018. Operations: lab interpreter ($interpret), care-gaps reminders ($care-gaps) with an embedded MCP-App view, and ChatGPT-connector search/fetch. Connectors: Fasten TEFCA, HealthEx, HBO, Flexpa, Epic, MEDENT, Open Wearables, SMART Health Links. Also a Claude Code plugin and OpenAI/Gemini adapters.

Try it in 60 seconds — no clone, no keys

The hosted demo runs synthetic data behind the full guardrail stack:

# Watch the deployment grade its own guardrails (PHI redaction, audit, step-up, ...):
curl "https://app.healthclaw.io/r6/fhir/\$conformance?format=text"

Point any MCP client at the public demo server — URL https://mcp-demo-production-ee2c.up.railway.app/mcp, no key required — then ask: "Search my health records for lab results and explain them in plain language." The demo server is unauthenticated but hard-pinned to a synthetic demo tenant, so it can only ever serve fake data. A separate production endpoint (mcp-server-production-5112) requires a deployment-scoped Authorization: Bearer <token> — real records stay behind auth, always. Hosted connectors cannot attach that header, so the demo URL above is the one to paste. One-command installs: gemini extensions install https://github.com/aks129/HealthClawGuardrails · claude plugin marketplace add aks129/HealthClawGuardrails · skills on ClawHub

Non-developer? Step-by-step guides for Claude (web/desktop/phone), Perplexity, ChatGPT, and Telegram — plus a 10-minute demo script — in docs/quickstarts/.

Listed in: Official MCP Registry (io.github.aks129/healthclaw-guardrails) · Glama (hosted connector) · ClawHub (14 skills) · Gemini CLI Extensions · agent-skills discovery at /.well-known/agent-skills/

Related MCP server: MCP Server for Google Cloud Healthcare API

Release highlights

Full notes live in Releases.

Version

Highlights

v1.10.0

Runs in front of a real FHIR server. The proxy now authenticates to an upstream FHIR server with its own client credential, so an agent never holds one — with a runnable Aidbox example that stands the guardrails in front of Aidbox and asserts each property rather than narrating it · access kernelr6.access becomes the one tenant reader, step-up gate, audit call and FHIR exit, adopted blueprint by blueprint · security: a caller-supplied seed bundle takes the ingest gate, not the mint gate (an unauthenticated write path, found and closed) · MCP: an expired session returns 404, so a client re-initializes instead of failing · demo data: multi-year synthetic blood-pressure history and a server-rendered trend chart, with home and clinic readings modelled distinctly · a defect catalogue wired into the PR gate, and drift guards that replay the published example's own claims against the running app

v1.9.0

CareAgents — the hosted consumer experience: sign up with a passkey, connect records through a pluggable connector marketplace (Fasten, Apple Health via Open Wearables, sample data), and spin up a guardrailed health agent reachable on web, Telegram, and iMessage · advisor registry — specialties ported from SmartHealthConnect (healthy-habits, care-completion, medication-refills, diet-exercise) as prompt-blocks over the guarded tool set, deferred ones honestly labeled · versioned informed consent enforced server-side (HTTP 428) before any real-record connection · forms rail ships end-to-end$populate → per-item human review (NKA never inferred) → provenance-stamped PDF → signed expiring link · error fidelity is conformance property seven (Grade A = 7/7), hardened across both MCP transports with a Python↔TypeScript drift guard · MCP Apps — care-gaps results embed an engine-served UI (text/html; profile=mcp-app) whose only fetch target is the guarded operation · security pass: fail-closed prod config, authenticated tenant reads, MCP transport auth, Alembic · SmartHealthConnect archived (skills frozen at v1.2.0; advisors are the live successors)

v1.8.0

Real-actions foundation — an agent can propose a real-world action (call, SMS, form) but commit only submits it (HTTP 202); execution happens through a separate approval that requires a single-use step-up credential and an expiry-guarded atomic claim, so the agent's own toolchain can never approve its own action (the spoofable X-Human-Confirmed header is gone) · ActionExecutor plugin registry — add a real-world capability behind the full guardrail rail in ~50 lines, no core changes (extend it) · mandatory red-flag emergency screen; fail-loud rails (no silent simulation) · durable execution — attempt ledger, provider reconciliation, external-tick reaper, append-only action-event log · reliability floor — config preflight (GET /r6/ops/preflight), Postgres CI lane, MCP fetch timeouts, poller 409-storm detection, source-aware resource identity (tenant, type, id), Fasten hardening + zombie-job reaper · public ROADMAP + contributor on-ramp · fixes: upstream FHIR error fidelity, quality measures default to current year

v1.7.0

Preventive care-gaps engine (Patient/$care-gaps, USPSTF/ACIP/ADA + eCQM crosswalk) · patient connect flow: identity-verified Fasten onboarding mints a webhook-gated, read-scoped 30-day agent token · prescription transfer requests (rx_transfer_request, Schedule II refused) — 29 MCP tools · per-agent quickstarts (Claude/Perplexity/ChatGPT/Telegram) · HBO export→FHIR converter + embedded-XML PHI scrubber · hardening: fail-closed webhook verify, scoped tokens, serverless write guard, live-path contract tests · clinical fixes: SNOMED diabetes detection, inclusive panic thresholds, one-sided-range honesty

v1.6.0

Lab reference-range interpreter (Observation/$interpret) · NQF 0018 quality measure (Measure/$evaluate-measure) · any-agent-framework adapters (OpenAI/Gemini) · Medplum-in-front recipe · SMBP triage on 2025 AHA/ACC · ruff lint gate · all dependency advisories remediated

v1.5.0

Read-auth hardening (tenant reads authenticated, not just scoped) · HL7 SDC forms — $populate / $extract

v1.4.0

Six health-data connectors (Fasten TEFCA, HealthEx, Health Bank One, Flexpa, Epic, MEDENT) behind one guardrail stack

v1.3.0

Wearables → FHIR Observations (8 providers, LOINC/UCUM mapping, device Provenance)

v1.2.0

Compiled Truth — current state + append-only Provenance trail per resource

What It Does

This is a vendor-neutral guardrail proxy that sits between any AI agent and any FHIR server. Every request passes through:

  • PHI redaction — Names truncated to initials, identifiers masked, addresses stripped, birth dates truncated to year

  • Immutable audit trail — Every read/write logged with tenant, agent, timestamp

  • Step-up authorization — HMAC-SHA256 tokens required for writes

  • Human-in-the-loop — Clinical writes blocked until a human confirms (HTTP 428); real-world actions (calls, SMS, forms) go further: commit only submits, and execution requires a provably out-of-band single-use approval the agent's own toolchain cannot satisfy

  • Tenant isolation — Every query scoped to tenant, cross-tenant access blocked

  • Medical disclaimers — Injected on all clinical resource reads

  • Compiled Truth — Current state + append-only evidence trail for every resource

AI Agent ──▶ MCP Server ──▶ Guardrail Proxy ──▶ Any FHIR Server
                              ↓                    (HAPI, Epic,
                         PHI redaction              Medplum, etc.)
                         Audit trail
                         Step-up auth
                         Human-in-the-loop

Prove it: guardrail conformance

The guardrails are verifiable, not marketing. A runnable harness probes any deployment with synthetic data and emits a scorecard across all seven properties — run it against your own instance (or ours):

python scripts/guardrail_conformance.py \
  --base-url https://app.healthclaw.io --tenant desktop-demo \
  --step-up-token "$(mint a token via POST /r6/fhir/internal/step-up-token)"
HealthClaw Guardrail Conformance — https://app.healthclaw.io [tenant=desktop-demo]
  Grade: A   (7/7 properties)
  [PASS] PHI Redaction            [PASS] Human-in-the-Loop
  [PASS] Immutable Audit Trail    [PASS] Tenant Isolation
  [PASS] Step-Up Authorization    [PASS] Medical Disclaimers
  [PASS] Error Fidelity — A (local-fhir-only)

Or hit the one-URL self-test on any running deployment — no token needed, it self-tenants internally and returns 200 at Grade A (503 otherwise):

curl "https://app.healthclaw.io/r6/fhir/\$conformance?format=text"

The local FHIR profile is Grade A: unsupported local-search inputs are rejected or reported according to Prefer: handling, and every failure path is audited. The same harness runs against the Flask test client as a CI baseline (tests/test_guardrail_conformance.py). --json emits a machine-readable report; --mcp-url additionally grades MCP tools/call error signaling as a separate profile. For an authenticated MCP deployment, set MCP_AUTH_TOKEN or pass --mcp-auth-token. Library API: from r6.conformance import LiveProbeClient, ProbeContext, run_conformance.

What this grade means (and what it doesn't)

The grade covers the HealthClaw guardrail layer only — a self-test of the seven properties against synthetic data it just created. It is not a HIPAA Security Rule assessment, a third-party audit, or a penetration test of your deployment: infrastructure, BAAs, encryption at rest/in transit, and access controls remain the deployer's responsibility (see Known Limitations). Because the harness is deployment-agnostic, a third party can run it against any instance as one input to a real assessment — it does not substitute for one. The report states this scope itself in every output format.

Install as a Claude Plugin

HealthClaw ships as a Claude Code plugin marketplace. Two plugins are available:

# Add the marketplace
claude plugin marketplace add aks129/HealthClawGuardrails

# Install the FHIR guardrail plugin (this repo)
claude plugin install healthclaw-guardrails@healthclaw-marketplace

# Install the personal-health companion plugin (frozen — upstream archived)
claude plugin install smarthealthconnect@healthclaw-marketplace

Plugin

Skills

Source

healthclaw-guardrails

curatr, fasten-connect, fhir-r6-guardrails, fhir-upstream-proxy, healthex-export, phi-redaction

aks129/HealthClawGuardrails

smarthealthconnect

care-completion, diet-exercise, healthy-habits, kids-health, medication-refills, research-monitor

aks129/SmartHealthConnect (archived — skills frozen at v1.2.0; live successors are CareAgents advisors)

Each skill is auto-discoverable — Claude loads it when your prompt matches the skill's trigger phrases (e.g. "check my care gaps", "redact this bundle", "run Curatr on my conditions").

Not on Claude/MCP? The same 28 guardrailed tools run on OpenAI, Gemini, LangChain, or plain HTTP via the framework-neutral bridge in adapters/ — see Recipe: run HealthClaw tools on any agent framework. Guardrails stay server-side, so no framework can bypass them.

Quick Start

# Install dependencies
uv sync

# Apply deterministic database migrations
STEP_UP_SECRET=your-secret uv run flask --app main init-db
STEP_UP_SECRET=your-secret uv run flask --app main seed-demo --tenant-id desktop-demo

# Run (local mode with SQLite)
STEP_UP_SECRET=your-secret python main.py

# Run with upstream FHIR server
FHIR_UPSTREAM_URL=https://hapi.fhir.org/baseR4 STEP_UP_SECRET=your-secret python main.py

# Open browser
open http://localhost:5000            # Landing page with live demo
open http://localhost:5000/r6-dashboard  # Interactive dashboard

Docker

docker-compose up -d --build

# macOS note: port 5000 conflicts with AirPlay Receiver — remap with:
# HOST_PORT=5050 docker-compose up -d --build

# Services:
# - fhir-mcp-guardrails (Flask, port 5000)
# - agent-orchestrator (MCP server, port 3001)
# - redis (port 6379)

MCP Tools (29)

Tool names use underscores (not dots) for Claude Desktop / MCP client compatibility.

Read tools (no step-up for public tenants):

Tool

Description

context_get

Retrieve pre-built context envelopes

fhir_read

Read a FHIR resource (redacted)

fhir_search

Search with patient, code, status, date filters

fhir_validate

Structural validation

fhir_stats

Observation statistics (count/min/max/mean)

fhir_lastn

Most recent N observations per code

fhir_interpret_labs

Lab reference-range interpretation ($interpret) — decision support, not diagnosis

care_gaps

Preventive-care gaps ($care-gaps) — screenings/immunizations that may be due, from the patient's own records

guardrail_conformance

Run the guardrail conformance self-test — graded A–F scorecard across all seven properties

fhir_permission_evaluate

R6 Permission access control evaluation

fhir_subscription_topics

List available SubscriptionTopics

questionnaire_populate

SDC $populate — pre-fill a Questionnaire for a subject

curatr_evaluate

Evaluate a FHIR resource for data quality issues

action_status

Poll a real-world action (call/SMS)

search

ChatGPT-connector-compatible search — thin wrapper over fhir_search, returns compact {id, title, url} results

fetch

ChatGPT-connector-compatible fetch by ResourceType/id — thin wrapper over fhir_read, returns {id, title, text, url, metadata}

Write tools (require step-up token):

Tool

Description

fhir_propose_write

Validate + preview without committing

fhir_commit_write

Commit with step-up auth + human-in-the-loop

questionnaire_extract

SDC $extract — extract resources from a completed QuestionnaireResponse

curatr_apply_fix

Apply patient-approved fixes with Provenance tracking

action_propose / action_commit

Propose / commit a real-world phone call or SMS

rx_transfer_request

Draft a pharmacy-transfer request call from active meds (Schedule II refused); commit via action_commit

shl_generate

Generate an encrypted SMART Health Link (QR)

Utility tools:

Tool

Description

fhir_get_token

Issue a 5-minute step-up token (call before any write)

fhir_seed

Seed a tenant with demo Patient + Observations + Condition

fhir_compiled_truth

Current state + Provenance evidence timeline

All tools add _mcp_summary with reasoning, clinical context, and limitations.

Guardrail Demo

The 6-step demo at /r6/fhir/demo/agent-loop shows the full guardrail sequence:

  1. PHI Redaction — Agent reads a patient, receives redacted data

  2. $validate Gate — Agent proposes an Observation, validated before write

  3. Permission Deny — No Permission rule exists, access denied with reasoning

  4. Permission Permit — Permit rule created, re-evaluation succeeds

  5. Step-up + Human-in-the-loop — Write requires both token and human confirmation

  6. Commit + Audit — Write succeeds, full audit trail generated

Comparison

Feature

This Project

AWS HealthLake MCP

Medplum MCP

Raw FHIR API

Works with any FHIR server

Yes

HealthLake only

Medplum only

N/A

PHI redaction on reads

Yes

No

No

No

Immutable audit trail

Yes

CloudTrail (separate)

Partial

No

Step-up auth for writes

Yes

IAM (separate)

Medplum auth

No

Human-in-the-loop

Yes

No

No

No

Permission $evaluate (R6)

Yes

No

No

No

Setup time

10 seconds

30+ minutes

15+ minutes

Varies

FHIR Version Support

Version

Profile

Status

Resources

R4

US Core v9

Stable

Patient, Condition, AllergyIntolerance, Immunization, MedicationRequest, Procedure, DiagnosticReport, CarePlan, CareTeam, Goal, DocumentReference, Coverage, ServiceRequest, Location, Organization, Practitioner, PractitionerRole, RelatedPerson, Specimen, FamilyMemberHistory

R6

v6.0.0-ballot3

Experimental

Permission, SubscriptionTopic, DeviceAlert, NutritionIntake, DeviceAssociation, NutritionProduct, Requirements, ActorDefinition

Both R4 and R6 resources flow through the same guardrail stack (PHI redaction, audit, step-up auth, tenant isolation). R6 ballot resources may change before final release.

Testing

# Python tests (1,490+ across 90+ files; includes action-rail, SDC, quality, labs, ops, CareAgents suites)
uv run python -m pytest tests/ -v
uv run python -m pytest tests/test_r6_routes.py::test_name -v   # single test

# MCP server tests
cd services/agent-orchestrator && npm ci && npm test

# Playwright end-to-end tests (UI + API, requires Flask on :5000)
cd e2e && npm ci && npx playwright install --with-deps chromium && npm test
cd e2e && npm run test:headed    # headed browser
cd e2e && npm run test:ui        # interactive UI mode

API Endpoints

Endpoint

Method

Description

/r6/fhir/metadata

GET

CapabilityStatement

/r6/fhir/health

GET

Liveness probe (reports upstream status)

/r6/fhir/{type}

POST

Create resource (requires step-up)

/r6/fhir/{type}

GET

Search resources

/r6/fhir/{type}/{id}

GET

Read resource (redacted)

/r6/fhir/{type}/{id}

PUT

Update resource (requires step-up + ETag)

/r6/fhir/{type}/$validate

POST

Validate resource

/r6/fhir/Questionnaire[/{id}]/$populate

POST

SDC — pre-fill a QuestionnaireResponse from a subject

/r6/fhir/QuestionnaireResponse/$extract

POST

SDC — extract a transaction Bundle (?dryRun=true to preview)

/r6/fhir/{type}/{id}/$deidentify

GET

Conservative de-identification preview (expert review required)

/r6/fhir/Observation/$stats

GET

Observation statistics

/r6/fhir/Observation/$lastn

GET

Most recent observations

/r6/fhir/Permission/$evaluate

POST

R6 access control evaluation

/r6/fhir/SubscriptionTopic/$list

GET

Subscription topic discovery

/r6/fhir/Bundle/$ingest-context

POST

Bundle ingestion + context envelope

/r6/fhir/context/{id}

GET

Retrieve context envelope

/r6/fhir/AuditEvent

GET

Search audit events

/r6/fhir/AuditEvent/$export

GET

Export audit trail (NDJSON/Bundle)

/r6/fhir/demo/agent-loop

POST

6-step guardrail demo

/r6/fhir/oauth/*

*

OAuth 2.1 + PKCE + SMART discovery

/r6/fhir/{type}/{id}/$curatr-evaluate

GET

Evaluate resource data quality (Curatr)

/r6/fhir/{type}/{id}/$curatr-apply-fix

POST

Apply patient-approved fixes with Provenance

Local search accepts the parameters advertised by /r6/fhir/metadata. Unknown parameters default to lenient handling (a bounded search.mode="outcome" warning); Prefer: handling=strict returns a 400 OperationOutcome. Unsupported modifiers and malformed supported values always return 400. _count=0 and _summary=count are count-only searches. Self links contain exactly the applied, URL-encoded parameters, and audit output never echoes submitted filter values or arbitrary parameter names.

Upstream Proxy

Connect to real FHIR servers while keeping all guardrails active:

FHIR_UPSTREAM_URL=https://hapi.fhir.org/baseR4 python main.py
  • Reads: Fetched from upstream, then redacted + audited + disclaimers added

  • Searches: Forwarded with all query params, results redacted per entry

  • Writes: Validated locally first, then forwarded with step-up auth check

  • URL rewriting: Upstream URLs never leak to clients

Tested with: HAPI FHIR R4/R5, SMART Health IT, Epic Sandbox.

Put the guardrails in front of your FHIR server — recipe for running the redaction + audit + step-up + human-in-the-loop stack in front of Medplum (the same pattern works for Aidbox, Google Cloud Healthcare, or any FHIR R4 server): docs/recipes/healthclaw-in-front-of-medplum.md. A repeatable integration test (tests/test_medplum_in_front.py) proves a Medplum-returned Patient comes back redacted + audited and writes are step-up gated before reaching Medplum.

Curatr — Patient-Owned Data Quality

Curatr is a patient-facing data quality skill that evaluates FHIR health records for coding issues and lets the patient decide how to resolve them.

1. Patient connects data → HealthClaw Guardrails deidentifies and loads it
2. OpenClaw calls curatr.evaluate → checks codes against live terminology APIs
3. Issues presented in plain language with impact and fix suggestions
4. Patient approves fixes → curatr.apply_fix updates resource + creates Provenance
5. Optional: generate a structured correction request for the source provider

What Curatr checks on a Condition:

Check

Service

Example

Deprecated code system

Local lookup (no network)

ICD-9-CM → critical

ICD-10-CM code validity

NLM Clinical Tables API

Invalid code → warning

SNOMED CT / LOINC validity

tx.fhir.org (HL7 public)

Unknown code → warning

RxNorm drug code

RXNAV API (NLM)

Missing RXCUI → warning

Display name accuracy

Cross-checked with canonical term

Mismatch → suggestion

Missing required fields

Structural

No clinicalStatus → warning

Every fix creates a linked Provenance resource recording patient intent, field changes, and agent attribution. All changes are audited in the immutable trail.

OpenClaw skill: skills/curatr/SKILL.md

Patient-controlled encrypted record sharing via QR code, implemented on top of jmandel/kill-the-clipboard-skill (MIT, pinned fa0020d) — credit Josh Mandel. HealthClaw governs what enters the bundle (step-up auth, profiles, guardrails, audit trail); KTC governs sharing (zero-knowledge server-side storage, SHL STU 1 protocol, revocation, in-browser viewer).

What it does: The shl_generate MCP tool (Write group, step-up required) fetches the patient's guardrailed FHIR bundle, encrypts it client-side in the MCP server (the SHL server never sees plaintext), uploads ciphertext, and returns:

  • shlink — the shlink:/ URI to encode in a QR (an encrypted pointer, not data)

  • viewer_link — browser URL for clinic staff

  • manage_link — patient-only revocation + access-log URL

Security: The QR encodes only the encrypted pointer. PHI never appears in the QR image. The SHL server stores only ciphertext + sha256(auth_token). Persona hard rule: see skills/share-health-qr/SKILL.md — never direct-encode PHI into QR images (incident 2026-06-12).

Quick Start (local)

# Start the SHL storage server (profile `shl`)
docker-compose --profile shl up -d

# Tell the MCP server where the SHL server lives
# Add to services/agent-orchestrator/.env or export:
export SHL_SERVER_URL=http://localhost:8000

Without SHL_SERVER_URL, shl_generate returns an explicit simulation stub (simulated: true) — never a fake link.

Railway Deploy

# 1. Add the SHL service
railway add --service shl-server

# 2. Attach a persistent volume (SQLite lives here)
railway service shl-server && railway volume add --mount-path /data

# 3. Configure the SHL server
railway variables --service shl-server \
  --set BASE_URL=<public-url-of-shl-server> \
  --set DB_PATH=/data/db.sqlite

# 4. Expose a public domain
railway domain --service shl-server

# 5. Deploy — MUST run from the shl-server directory
cd services/shl-server && railway up --service shl-server

# 6. Wire the MCP server to the SHL server
railway variables --service mcp-server \
  --set SHL_SERVER_URL=<public-url-of-shl-server>

Caveat 1 — deploy from the right directory: The repo-root railway.toml targets the Flask Dockerfile. If you run railway up --service shl-server from the repo root, Railway uses the wrong Dockerfile and the deploy fails. Always cd services/shl-server first — that directory has its own railway.toml that points to the correct image.

Caveat 2 — watchPatterns skip: A service that inherited watchPatterns from the root config may silently skip Dockerfile-only deploys (no source file changes detected). The per-service railway.toml in services/shl-server/ overrides this after the first successful build. If deploys are skipped, force one with railway up --service shl-server from the shl-server directory.

Caveat 3 — simulation mode: Without SHL_SERVER_URL on the MCP server, shl_generate returns { simulated: true, note: "SHL_SERVER_URL not configured — returned stub." }. Personas surface this note verbatim and never improvise an alternative.

OpenClaw skill: skills/share-health-qr/SKILL.md

R6-Specific Resources (Experimental)

These resources are part of the FHIR R6 ballot3 specification and may change before final release.

Resource

What's New in R6

Permission

Access control (separate from Consent), $evaluate operation

SubscriptionTopic

Restructured pub/sub (introduced R5, maturing R6)

DeviceAlert

ISO/IEEE 11073 device alarms

NutritionIntake

Dietary consumption tracking

DeviceAssociation

Device-patient relationships

NutritionProduct

Nutritional product definitions

Requirements

Functional requirements tracking

ActorDefinition

Actor role definitions

US Core v9 R4 Resources (Stable)

Standard FHIR R4 resources conforming to US Core Implementation Guide v9. These are widely deployed in US healthcare and stable for production use.

AllergyIntolerance, Immunization, MedicationRequest, Medication, MedicationDispense, Procedure, DiagnosticReport, CarePlan, CareTeam, Goal, DocumentReference, Location, Organization, Practitioner, PractitionerRole, RelatedPerson, Coverage, ServiceRequest, Specimen, FamilyMemberHistory

Environment Variables

Variable

Required

Default

Description

STEP_UP_SECRET

Production

HMAC-SHA256 signing secret

FHIR_UPSTREAM_URL

No

Upstream FHIR server (enables proxy mode)

SQLALCHEMY_DATABASE_URI

Production

sqlite:///mcp_server.db

Database connection

SESSION_SECRET

No

(dev key)

Flask session secret

READ_AUTH_ENABLED

Production

false

Require tenant-bound credentials on protected reads

PUBLIC_TENANTS

Production

Explicit comma-separated synthetic/demo tenant allowlist

REDIS_URL

Production

Shared nonce, OAuth, rate-limit, and worker state

MCP_AUTH_TOKEN

HTTP MCP

Bearer credential required by MCP HTTP transports

MCP_PUBLIC_DEMO

No

false

Run an unauthenticated MCP server hard-pinned to a synthetic demo tenant (the public keyless demo). Never set on a server that reaches real tenants

MCP_DEMO_TENANT

No

desktop-demo

Synthetic tenant the demo server is pinned to when MCP_PUBLIC_DEMO is set

FHIR_UPSTREAM_TIMEOUT

No

15

Upstream request timeout (seconds)

FHIR_LOCAL_BASE_URL

No

Local URL for response URL rewriting

Database DDL is never run during WSGI import. Run flask --app main init-db before each release; it applies the locked Alembic revisions. Operators adopting Alembic on an existing v1.8.0 Postgres deployment must follow the database migration runbook to verify and stamp the compatibility baseline before upgrading.

Project Structure

main.py                         Flask app entry point
app.py                          Web UI routes (landing, dashboard)
r6/
  routes.py                     R6 FHIR REST Blueprint (1,732 lines)
  models.py                     R6Resource, ContextEnvelope, AuditEventRecord
  validator.py                  FHIR R6 structural validation
  redaction.py                  PHI redaction (names, identifiers, addresses, DOB, telecom)
  audit.py                      Immutable AuditEvent recording
  stepup.py                     HMAC-SHA256 step-up token management
  oauth.py                      OAuth 2.1 + PKCE + SMART-on-FHIR discovery
  health_compliance.py          Disclaimers, HITL, de-identification preview, audit export
  context_builder.py            Bundle ingestion + context envelopes
  rate_limit.py                 Per-tenant rate limiting
  fhir_proxy.py                 Upstream FHIR server proxy with URL rewriting
  curatr.py                     Curatr data quality engine (terminology lookups + fix application)
services/agent-orchestrator/
  src/index.ts                  MCP server (Streamable HTTP + SSE)
  src/tools.ts                  12 tool definitions + executor (incl. curatr.evaluate, curatr.apply_fix)
e2e/                            Playwright end-to-end tests
templates/                      Jinja2 (landing page, dashboard)
static/                         CSS + JS for interactive dashboard
skills/curatr/                  Curatr OpenClaw skill definition
tests/                          266 pytest tests (8 files, incl. test_us_core_r4.py)

Personal FHIR data store — patient import flow

This walkthrough shows how to go from a raw HealthEx export to querying your own records through Claude Code's MCP tools.

1. Start the stack

uv sync
uv run python main.py                         # Flask on :5000
cd services/agent-orchestrator && npm ci && npm start  # MCP on :3001

2. Import your HealthEx / Flexpa / generic FHIR bundle

# Dry-run first to preview without writing
python scripts/import_healthex.py \
  --bundle-file ~/Downloads/my-records.json \
  --dry-run

# Real import — prints context_id on success
python scripts/import_healthex.py \
  --bundle-file ~/Downloads/my-records.json \
  --tenant-id my-patient \
  --step-up-secret "$STEP_UP_SECRET"

3. Connect Claude Code via MCP

.mcp.json in this repo auto-configures Claude Code when you open the project. Update X-Tenant-ID to match your --tenant-id:

{
  "mcpServers": {
    "healthclaw-local": {
      "type": "http",
      "url": "http://localhost:3001/mcp",
      "headers": { "X-Tenant-ID": "my-patient" }
    }
  }
}

Then in Claude Code:

Use fhir_search to find all my Conditions
Use context_get with context_id <ctx-id> to get my full context envelope
Use curatr_evaluate on Condition/<id> to check data quality

4. Set up Fasten Connect (optional)

# .env additions
FASTEN_PUBLIC_KEY=<key>
FASTEN_PRIVATE_KEY=<key>
FASTEN_WEBHOOK_SECRET=<secret>
FASTEN_CURATR_SCAN=true    # auto-run Curatr after each import

Records arrive via webhook at /r6/fasten/webhook and are stored under the patient's canonical tenant ID.

5. Deidentify for sharing

# De-identification preview (not a legal Safe Harbor determination)
curl -H "X-Tenant-ID: my-patient" \
  http://localhost:5000/r6/fhir/Patient/pt-1/\$deidentify

# Patient-controlled (preserves birthDate, strips institutional identifiers)
curl -H "X-Tenant-ID: my-patient" \
  "http://localhost:5000/r6/fhir/Patient/pt-1/\$deidentify?mode=patient-controlled&patient_id=my-patient"

6. Telegram bot (optional)

TELEGRAM_BOT_TOKEN=<token> TENANT_ID=my-patient \
FHIR_BASE_URL=http://localhost:5000/r6/fhir \
python openclaw/bot.py

Commands: /health, /conditions, /labs, /curatr, /curatr fix, /approve.

Or via Docker Compose:

docker-compose --profile openclaw up -d

7. Use Medplum as the backing FHIR store (optional)

Set in .env (leave FHIR_UPSTREAM_URL empty):

MEDPLUM_BASE_URL=https://api.medplum.com/fhir/R4
MEDPLUM_CLIENT_ID=<id>
MEDPLUM_CLIENT_SECRET=<secret>

All guardrails apply to Medplum responses identically to local SQLite mode. Access tokens are cached in Redis (key medplum:access_token; falls back to in-process cache when Redis is unavailable).


Known Limitations

  • The conformance grade is a self-test of the guardrail layer, not a HIPAA assessment or third-party audit — see What this grade means

  • Local mode: JSON blob storage with table-scan search (no indexed fields)

  • Redaction is HIPAA Safe-Harbor-style field redaction (demographics), not Expert Determination. It's a compensating control that removes identifier-class fields; it is not a legal de-identification determination. Production de-id rigor (profile-specific recursive allowlists, an Expert-Determination path) is on the roadmap (#112).

  • Validation is structural, not full StructureDefinition/profile conformance or terminology binding. What's demonstrated is the guardrail contract (redact + audit + step-up + human-confirm + tenant isolation + error fidelity), not production validation depth — that's tracked in #112.

  • SubscriptionTopic stored but notifications not dispatched

  • Clinical FHIR writes gate human-in-the-loop with a header flag (X-Human-Confirmed), not cryptographic confirmation — a compensating control for the demo, not proof a human acted. Real-world actions (phone/SMS/etc.) no longer use that header: commit only submits the action for out-of-band approval (202 awaiting_confirmation), and the patient's Approve tap consumes a single-use ActionConfirmation credential server-side before anything executes.

  • OAuth endpoints are for discovery/SMART advertisement; route enforcement is via step-up + read-auth tokens, and the auto-approve authorize flow is limited to public/demo tenants (no per-user consent screen)

  • No historical versioning (version_id increments but old versions not retrievable)

  • Upstream proxy: no response caching, no cross-version translation

  • Security is config-dependent — production requires READ_AUTH_ENABLED=true (authenticate non-public reads), INTERNAL_TOKEN_MINT_SECRET (gate token mint/seed for non-public tenants; fail-closed in prod when unset), PUBLIC_TENANTS limited to synthetic demo tenants, a real SESSION_SECRET/STEP_UP_SECRET, and https-only upstreams

  • Step-up tokens are valid for multiple writes within their 5-min TTL (not single-use); irreversible actions rely on state-machine idempotency (guarded WHERE status='proposed' claim) rather than nonce consumption

Contributing — this is a community effort

HealthClaw Guardrails is developed in the open as a shared reference, not a commercial product. The guardrail layer between AI agents and clinical data only gets trustworthy if a lot of people with different vantage points pressure-test it. We especially want:

  • Implementers building FHIR × MCP integrations — tell us where the patterns break in the real world.

  • Clinicians & compliance folks — challenge the redaction profiles, audit model, and the documented HIPAA postures.

  • Standards people (HL7 / SDC / SMART) — tell us where we've diverged from the spec, especially on $populate/$extract.

  • Anyone — open an issue, file a "you got this wrong," or send a PR.

Start here: CONTRIBUTING.md · Roadmap · Dev Guide · Code of Conduct · CHANGELOG.md · Security policy

Good first contributions are labeled in the issue tracker. Contributions are DCO-signed (git commit -s) under the MIT license — see LICENSING.md for the project's licensing posture going forward.

Community

License

MIT — free to use, fork, and build on. See LICENSE.

Available Tools

29 tools
action_commitSubmit Real-World Action for ConfirmationA
Destructive
Inspect

Submit a previously proposed action for the patient's OWN out-of-band confirmation (their dashboard or Telegram) AFTER they've reviewed and verbally/textually agreed to the draft. Requires step-up authorization (call fhir_get_token first; pass as _stepUpToken). This call does NOT execute anything and never accepts or sends any 'human confirmed' flag — only the patient tapping Approve in their own out-of-band channel can trigger execution. Returns status 'awaiting_confirmation' and is terminal for your turn: do not call action_commit again for the same action_id. Use action_status to check whether the patient has approved yet.

ParametersJSON Schema
NameRequiredDescriptionDefault
action_idYesID returned by action_propose

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds crucial context beyond annotations: it clarifies the call does not execute anything, never accepts a 'human confirmed' flag, and is terminal. This aligns with openWorldHint and adds detail not present in destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise at 4 sentences, each sentence adds unique value. Front-loaded with main purpose. Could trim minor redundancy but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, description explains return status and terminal behavior. Includes prerequisite step-up token. Lacks only minor details like error cases, but sufficient for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter action_id, so description adds no new meaning. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool submits a proposed action for out-of-band patient confirmation, clearly distinguishing it from siblings like action_propose and action_status. It uses specific verbs and resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when to use (after patient agrees to draft) and when not (do not call again for same action_id). Mentions prerequisites (fhir_get_token) and alternatives (action_status).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

action_proposePropose Real-World ActionA
Read-only
Inspect

Propose a real-world action (phone call or SMS) on the patient's behalf. Returns a draft (id + script) the patient MUST review before submitting via action_commit. Does not execute anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesAction type
payloadYesAction content: { to: recipient label, phone: number to dial/text, body: call script or message text }

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate read-only and non-destructive behavior; the description confirms no execution and adds the draft review step, providing full transparency without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load purpose and behavior, with zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description adequately explains the return value (draft with id + script) and the required workflow, making it complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with detailed descriptions for both parameters; the tool description adds no new information beyond what the schema already provides, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool proposes a real-world action (phone call or SMS) on the patient's behalf and distinguishes it from the sibling action_commit by noting it returns a draft for review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the draft must be reviewed before submitting via action_commit and clarifies that it does not execute anything, providing clear when-to-use and alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

action_statusAction StatusA
Read-only
Inspect

Check the status and outcome of an action (proposed/awaiting_confirmation/executing/completed/failed/needs_review/unknown/expired). needs_review means it ran but the outcome could not be confirmed - show the patient the evidence. unknown means the provider MAY have acted - never re-propose the same action. Use after action_commit to see whether the patient has approved yet, and to report the final result back to them.

ParametersJSON Schema
NameRequiredDescriptionDefault
action_idYesID returned by action_propose

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the tool is known to be safe. The description adds valuable behavioral context by interpreting statuses (e.g., 'needs_review means it ran but the outcome could not be confirmed', 'unknown means the provider MAY have acted - never re-propose the same action'), going beyond what annotations provide. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences covering purpose, status list, and usage guidance. Front-loaded with key information. No redundant or irrelevant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one fully documented parameter, no output schema, and no nested objects, the description provides all necessary context: what the tool does, when to use it, and how to interpret results. Status semantics are explained clearly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'action_id' has schema description 'ID returned by action_propose'. With 100% schema coverage, the description adds minimal additional value but does clarify the source of the ID, which is helpful context. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Check the status and outcome of an action' with a specific list of status values (proposed, awaiting_confirmation, etc.), clearly defining the tool's purpose and distinguishing it from sibling tools like action_commit and action_propose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Use after action_commit to see whether the patient has approved yet' and provides conditional guidance for 'needs_review' and 'unknown' statuses, including what actions to take (show evidence, never re-propose).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

care_gapsPreventive Care GapsA
Read-only
Inspect

Check which preventive-care screenings/immunizations a patient may be due for (blood pressure, cholesterol, colorectal/cervical/breast cancer screening, flu, diabetes A1c), from their own connected records. Decision support based on USPSTF/ACIP/ADA guidelines — not a diagnosis or directive.

ParametersJSON Schema
NameRequiredDescriptionDefault
subjectNoPatient reference (e.g. 'Patient/pt-1')

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark it as read-only and non-destructive. The description adds that it uses the patient's own connected records and is decision support (not diagnosis), clarifying its scope and limitations beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence followed by a clarifying caveat. Every word adds value, no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity and lack of output schema, the description adequately explains what the tool does and its basis. However, it omits the return format, which would improve completeness. Still, it's sufficient for an AI.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes the sole parameter 'subject' with an example. The description mentions 'patient' but adds no further semantics beyond the schema. With 100% schema coverage, a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check') and resource ('preventive-care screenings/immunizations') and lists concrete examples. It clearly distinguishes from sibling tools by focusing on preventive care gaps rather than generic FHIR operations or data fetching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use to check due screenings/immunizations based on guidelines. It does not explicitly state when not to use it or name alternatives, but the context and sibling list help. A moderate gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_getGet Health ContextB
Read-only
Inspect

Retrieve a pre-built context envelope with patient-centric FHIR resources. Returns bounded, policy-stamped, time-limited context.

ParametersJSON Schema
NameRequiredDescriptionDefault
context_idYesContext envelope ID

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety profile (readOnlyHint, destructiveHint). Description adds context about bounded, policy-stamped, time-limited return, which goes beyond annotations but is not extensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, each adding value. Front-loaded with action and key properties, no redundant or unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given simple tool with one parameter and annotations, description adequately explains what is returned and key attributes. Could elaborate on 'context envelope' for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameter meaning is clear from schema. Description adds no specific parameter details beyond the schema, achieving baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it retrieves a context envelope with FHIR resources, distinguishing it from generic read tools. However, it does not explicitly contrast with sibling tools like fhir_read or fhir_search, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no prerequisites or exclusions provided. Agent must infer usage from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

curatr_apply_fixApply Data Quality FixAInspect

Apply patient-approved data quality fixes to a FHIR resource. Creates a linked Provenance record with full attribution. Requires step-up authorization (X-Step-Up-Token) and human confirmation (X-Human-Confirmed: true) for clinical resources like Condition.

ParametersJSON Schema
NameRequiredDescriptionDefault
fixesYesList of field fixes to apply. Each fix has 'field_path' (dot-notation, e.g. 'Condition.code.coding[0].system') and 'new_value' (the corrected value).
resource_idYesID of the resource to fix
resource_typeYesFHIR resource type to fix (e.g. 'Condition')
patient_intentYesPlain-language reason for the fix, provided by the patient (recorded in Provenance).

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate the tool is not read-only and not destructive. The description adds that it applies fixes and creates a Provenance record, plus authorization requirements. This adds useful behavioral context beyond the annotations, though it could detail consequences (e.g., whether the original resource is versioned).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences and a header line concisely convey purpose, requirements, and core functionality. No filler; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains what the tool does and its prerequisites, but it does not mention the return value or outcome (e.g., updated resource, success status). Given no output schema, this is a gap. However, the level of detail is adequate for a mutation tool with good parameter documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are documented. The tool description briefly restates the 'fixes' array structure and patient_intent purpose, adding minimal new meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool applies patient-approved data quality fixes to FHIR resources and creates a Provenance record. It uses specific verbs and resources, but does not explicitly differentiate from sibling tools like curatr_evaluate or fhir_commit_write, though the patient-approval context provides some distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions prerequisites: step-up authorization and human confirmation headers for clinical resources. This provides clear context on when the tool is appropriate, though it does not explicitly state when not to use it or list alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

curatr_evaluateEvaluate Data QualityA
Read-only
Inspect

Evaluate a FHIR resource for data quality issues. Checks coding elements against public terminology services (tx.fhir.org for SNOMED/LOINC, NLM for ICD-10-CM, RXNAV for RxNorm) and structural rules. Returns issues in plain language with patient-facing impact descriptions and resolution suggestions. Read-only — no step-up required.

ParametersJSON Schema
NameRequiredDescriptionDefault
resource_idYesID of the resource to evaluate
resource_typeYesFHIR resource type to evaluate (e.g. 'Condition')

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond the annotations: it specifies external terminology services used (tx.fhir.org, NLM, RXNAV), output format (plain language with impact descriptions and suggestions), and confirms read-only with 'no step-up required.' No contradiction with readOnlyHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences, front-loaded with the core purpose, and each sentence adds distinct value: purpose, technical detail, and behavioral trait. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (evaluating FHIR data quality with multiple services), the description covers the key aspects: what it does, how it does it, what output looks like, and its read-only nature. Even without an output schema, the output description is sufficient. Annotations cover safety profile. The description feels complete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters. The description does not add new information about parameter semantics beyond what the schema already provides, thus baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates a FHIR resource for data quality issues, specifying it checks coding elements against public terminology services and structural rules. This distinguishes it from sibling tools like fhir_validate or guardrail_conformance, which are more about validation or conformance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for data quality evaluation of FHIR resources, but does not explicitly state when to use this tool over alternatives like fhir_validate or curatr_apply_fix. It provides helpful context about the checks performed, but lacks direct guidance on exclusions or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchFetch Health RecordA
Read-only
Inspect

ChatGPT-connector-compatible fetch of one FHIR resource by id ('ResourceType/id', as returned by search). Returns the full document (PHI-redacted server-side) with metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesResource reference: 'ResourceType/id'

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable behavioral detail: server-side PHI redaction, return of full document with metadata, and 'ChatGPT-connector-compatible' operation. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that packs essential information: purpose, id format, return content, and redaction. It is front-loaded and efficient, though it could benefit from slight restructuring for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple fetch tool with one parameter and no output schema, the description adequately covers the inputs, return value, and server-side processing. It does not discuss error handling or permissions, but the annotations cover safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the parameter description in the schema is identical to the usage in the tool description. The description adds no new meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it fetches a single FHIR resource by ID and returns the full document with PHI redacted. It specifies the id format 'ResourceType/id', distinguishing it from search endpoints. However, it does not explicitly differentiate from the sibling tool 'fhir_read', which likely serves a similar purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after a search by stating 'as returned by search', but it lacks explicit guidance on when to use this tool versus alternatives like 'fhir_read' or 'fhir_search'. No mention of when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_commit_writeCommit FHIR WriteC
Destructive
Inspect

Commit a previously proposed write. Requires step-up authorization token. This is a destructive operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYesThe FHIR resource to commit
operationYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already include destructiveHint: true, and the description redundantly states 'This is a destructive operation.' However, it adds value by disclosing the need for a step-up authorization token, which is not covered by annotations. The description goes beyond annotations but only marginally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with only two sentences, front-loading the core purpose. Every sentence contributes information. However, it could be slightly reordered for better impact, but overall it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, and the siblings include fhir_propose_write, the description should explain the relationship with the propose step (e.g., 'Call after fhir_propose_write to finalize'). It also lacks details about what gets destroyed or the return value. The description is incomplete for an agent to use correctly in context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not discuss any parameters. The input schema has two parameters with descriptions, but the 'operation' parameter description incorrectly repeats the resource description, reducing its usefulness. With 50% schema coverage and no compensatory information in the description, the parameter semantics are poorly supported.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Commit a previously proposed write', which clearly indicates the verb (commit) and the resource (previously proposed write). However, it does not explicitly differentiate from sibling tools like action_commit, so there's room for improvement in distinguishing from similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions a prerequisite ('Requires step-up authorization token') but provides no guidance on when to use this tool versus alternatives, nor does it exclude any inappropriate use cases. No when-not or alternative references are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_compiled_truthCompiled Truth TimelineA
Read-only
Inspect

Return the current best understanding of a FHIR resource plus the append-only evidence trail (Provenance entries) of how it got there. Use this before presenting resource-specific facts to a patient — surfaces curation_state and quality_score so the agent can say not just WHAT the record says but WHY it says it. Redacted, audited. Response includes _meta.ui.resourceUri pointing to an embeddable review UI.

ParametersJSON Schema
NameRequiredDescriptionDefault
resource_idYesID of the resource
resource_typeYesFHIR resource type (e.g. 'Condition', 'AllergyIntolerance')

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. Description adds 'Redacted, audited' and mentions curation_state and quality_score surfaces, plus an embeddable review UI link. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each adding value: first states core function, second gives usage context, third lists additional outputs. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has two required params and no output schema, but description adequately explains the key outputs (provenance, curation_state, quality_score, UI link). Could mention any rate limits or performance characteristics, but not essential for a read-only tool with annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers both parameters fully (100% coverage). Description does not add new details about parameters themselves, but explains what the tool produces (provenance, curation_state, quality_score, UI link), which indirectly clarifies the expected input usage. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it returns 'current best understanding of a FHIR resource plus the append-only evidence trail (Provenance entries)'. Distinguishes from siblings like fhir_read by focusing on compiled truth with provenance and quality indicators.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this before presenting resource-specific facts to a patient', providing specific context for its use. Does not mention when not to use or alternatives, but the sibling list is large and this guidance helps narrow down.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_get_tokenMint Step-Up TokenA
Read-only
Inspect

Get a fresh step-up authorization token for write operations. Call this before fhir_propose_write, fhir_commit_write, or curatr_apply_fix. Tokens expire after 5 minutes. Returns the token string — pass it as _stepUpToken in subsequent write tool calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
tenant_idYesTenant ID to scope the token to

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses token expiration (5 minutes) and that this is a pre-requisite for write operations. Annotations already indicate readOnlyHint=true, so the description adds value by explaining the token's purpose and lifecycle without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no fluff. Each sentence serves a purpose: first states what the tool does, second gives usage context and output detail. Perfectly front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool that returns a token string, the description covers purpose, when to use, return value, and how to use it. No output schema exists, but the description fully explains the return value usage. Complete and self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with one parameter (tenant_id) already described. The description does not add additional semantics beyond the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Get a fresh step-up authorization token for write operations', specifying the exact verb, resource, and purpose. It also names sibling tools that require this token, distinguishing it from other tools in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Call this before fhir_propose_write, fhir_commit_write, or curatr_apply_fix', providing clear when-to-use guidance. Also mentions token expiration (5 minutes) and instructs how to pass the token as _stepUpToken in subsequent calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_interpret_labsInterpret Lab ResultsA
Read-only
Inspect

Interpret lab Observations against reference ranges — flags each value low/normal/high/critical (HL7 v3 ObservationInterpretation) and returns clinician + consumer summaries. Decision support, not diagnosis. Read-tier.

ParametersJSON Schema
NameRequiredDescriptionDefault
bundleNoA FHIR Bundle of Observations to interpret
subjectNoPatient reference (e.g. 'Patient/pt-1') — interpret the tenant's stored Observations for this subject
observationNoA single FHIR Observation to interpret

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=true, destructiveHint=false) already indicate a safe read operation. The description adds context about decision support not being diagnosis, but does not disclose further behavioral traits such as error handling, authentication needs, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, delivering the core purpose and caveat in just two sentences. It is front-loaded and contains no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the description explains output (flags and summaries) and the read-tier nature, it does not address what happens when no parameters are provided or describe the response format in detail. Given no output schema and optional parameters, more context would be beneficial, but the description is adequate for common use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all three parameters. The tool description does not add significant meaning beyond the schema; it only implies that the tool can be called with a single observation or a bundle. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool interprets lab Observations against reference ranges, flags values, and returns summaries. It uses a specific verb ('Interpret') and resource ('lab Observations'), and distinguishes itself from sibling tools such as fhir_read or fhir_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes 'Decision support, not diagnosis. Read-tier,' which provides context on when to use the tool (decision support) and its read-only nature. However, it does not explicitly mention when not to use it or directly compare it to alternative tools for similar tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_lastnLatest ObservationsA
Read-only
Inspect

Get the last N observations per code. Standard FHIR $lastn (since R4). Returns most recent observations by storage order.

ParametersJSON Schema
NameRequiredDescriptionDefault
maxNoMax observations per code (default 1)
codeNoLOINC code filter
patientNoPatient reference filter

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, and the description aligns with 'Get' and 'Returns'. It adds behavior detail 'by storage order', which is useful. No contradictions. With annotations covering safety, the description provides additional ordering context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences. The first sentence efficiently states the core function, and the second adds standard reference and ordering behavior. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 optional parameters and no output schema, the description explains the core operation (last N per code, FHIR standard, storage order). It does not detail return format or empty results, but the information provided is sufficient for basic usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear parameter descriptions. The description only reinforces 'per code' for the code parameter but adds no extra meaning beyond the schema. Baseline score is appropriate as no additional semantic value is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get', resource 'observations per code', and explicitly references the standard FHIR $lastn operation. It distinguishes itself from sibling tools like fhir_search by specifying 'per code' and 'last N' semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates it's the standard way to retrieve last N observations per code but does not explicitly state when to use this tool versus alternatives like fhir_search. It lacks guidance on exclusions or prerequisites, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_permission_evaluateEvaluate Access PermissionA
Read-only
Inspect

Evaluate R6 Permission resources for access control decisions. Returns permit/deny based on stored Permission rules. Separates access control (Permission) from consent records (Consent).

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesAction to evaluate
subjectNoSubject reference (e.g., 'Practitioner/dr-1')
resourceNoResource reference to evaluate access for

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false, so the description does not need to repeat that. It adds that the tool returns permit/deny, which is behavioral, but no additional context on authentication, rate limits, or side effects. Given the annotations, the description is adequate but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at three sentences with no redundant information. It front-loads the purpose and adds a clarifying statement about the distinction from Consent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the description covers the core functionality (evaluate permission, return permit/deny) and clarifies the separation from consent. It lacks mention of edge cases or default behavior but is sufficient for a straightforward evaluation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage, so baseline is 3. The description does not add any extra meaning or context for the parameters beyond what the schema already provides (action enum, subject and resource strings).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates R6 Permission resources for access control decisions and returns permit/deny. It specifically mentions the resource type (R6 Permission) and the output, and distinguishes from Consent records. No sibling tool performs this exact function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some context by separating Permission from Consent, implying when to use this tool (for access control) vs. a consent-related tool. However, it does not explicitly state when to use or not use this tool, nor does it name any alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_propose_writePropose FHIR WriteA
Read-only
Inspect

Propose a write — validates the resource and returns a preview. Does NOT commit. Safe to call without step-up authorization.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYesThe FHIR resource to write
operationYesWrite operation type

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and non-destructive behavior. Description adds clarity that no commit occurs and no special authorization is needed, complementing annotations well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no fluff, front-loaded with purpose and key safety information. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Sufficient for understanding when to call, but lacks details about the preview output format. Minor gap given no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters. The description does not add additional meaning beyond what the schema provides, so baseline score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool validates a FHIR resource and returns a preview, explicitly noting it does not commit. Differentiates from sibling commit and validation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use for validation/preview before commit, but lacks explicit when-not-to-use or alternative naming. The 'Safe to call without step-up authorization' provides context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_readRead FHIR ResourceA
Read-only
Inspect

Read a specific FHIR resource by type and ID. Supports FHIR R4 US Core v9 stable resources and FHIR R6 ballot3 experimental resources. Returns redacted resource with PHI protection.

ParametersJSON Schema
NameRequiredDescriptionDefault
resource_idYesThe resource ID
resource_typeYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint and destructiveHint. The description adds value by stating the resource is redacted with PHI protection and supports specific FHIR versions, providing useful behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences: first states purpose and parameters, second adds supported versions and return behavior. No filler words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with no output schema, the description covers the key points: what it does (read specific resource), parameters (type and ID), constraints (supported versions, PHI redaction). Adequate for agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% (resource_id has minimal description, resource_type only enum). The description adds context that it reads 'by type and ID' but does not add parameter-specific details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads a specific FHIR resource by type and ID. It distinguishes from sibling tools like fhir_search (search) and fhir_commit_write (write).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for reading a single resource by ID, but lacks explicit when-to-use, when-not-to-use, or alternatives. It only mentions supported versions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_seedSeed Demo DataAInspect

Seed a tenant with a realistic Patient + Observations + Condition bundle for live testing. Use this at the start of a demo session to populate data. Returns created resource IDs and a ready-to-use step_up_token.

ParametersJSON Schema
NameRequiredDescriptionDefault
tenant_idNoTenant to seed (default: desktop-demo)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds that the tool writes (seeds) data and returns a step_up_token, which is useful context. No contradictions; the behavioral summary is transparent for a population tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no fluff. The purpose and usage are front-loaded, and every word adds value. Highly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one optional parameter and no output schema, the description covers the main points: action, when to use, and what is returned (IDs and token). Could specify format or more detail, but sufficient for a simple seeding tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Parameter schema has 100% coverage with a description for tenant_id. The tool description does not add extra meaning beyond the schema, which is adequate. Baseline 3 is appropriate as schema already documents the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (seed), specific resources (Patient + Observations + Condition bundle), and purpose (live testing, demo session). The description is distinct from sibling tools, which focus on reading, searching, or committing, not seeding demo data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this at the start of a demo session to populate data,' providing clear context. While it doesn't explicitly state when not to use, the demo/testing context is clear and implies production avoidance. No alternatives mentioned, but the tool is unique among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_statsObservation StatisticsA
Read-only
Inspect

Compute statistics (count, min, max, mean) over numeric Observation values. Standard FHIR $stats (since R4). Only supports valueQuantity. Filter by patient and/or code.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNoLOINC code to filter Observations (e.g., '2339-0' for Glucose)
patientNoPatient reference filter (e.g., 'Patient/pt-1')

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. The description adds that it is a standard FHIR operation and only supports valueQuantity, which is behavioral context beyond annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with three sentences. The first sentence immediately states the tool's purpose, and no extraneous information is included. Efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and no output schema, the description covers the key aspects: purpose, supported value type, and filtering. It does not detail the output format, but the listed statistics (count, min, max, mean) provide reasonable expectation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are described in the schema. The description adds meaning by explaining filtering and providing an example format (LOINC code) for the 'code' parameter, which goes beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool computes statistics (count, min, max, mean) over numeric Observation values using valueQuantity, and mentions FHIR R4 standard. This is specific and distinguishes it from sibling tools like fhir_search or fhir_interpret_labs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on filtering by patient and/or code, and specifies that only valueQuantity is supported. It implies usage for numeric observation statistics but does not explicitly exclude alternative scenarios or mention when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_subscription_topicsList Subscription TopicsA
Read-only
Inspect

List available SubscriptionTopics for event-driven subscriptions. R6 moves topic-based subscriptions toward Normative. Agents discover what events they can subscribe to.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description does not need to reiterate safety. It adds some context about R6 and normative status but does not disclose additional behavioral traits beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, concise and front-loaded with the main action. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description is complete. It explains what the tool does and why it is relevant (R6 normative status), which is sufficient for an agent to understand its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters, so schema description coverage is 100% by default. With zero parameters, the description does not need to add parameter semantics, and the baseline score of 4 is appropriate as there is no missing information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List' and the resource 'SubscriptionTopics', and distinguishes from siblings by specifying 'event-driven subscriptions' and 'discover what events they can subscribe to', which is not the purpose of other tools like fhir_read or fhir_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates when to use the tool: to discover events for subscriptions. It does not explicitly state when not to use it or provide alternatives, but the context is clear enough for agents to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fhir_validateValidate FHIR ResourceA
Read-only
Inspect

Validate a proposed FHIR R6 resource against structural rules. Returns OperationOutcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
resourceYesThe FHIR resource to validate

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, destructiveHint=false. Description adds that validation is structural and returns OperationOutcome. Does not contradict annotations; could elaborate on what happens on failure or whether it interacts with external systems.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences with front-loaded action. No redundant information; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single parameter with full schema coverage and annotations present, the description is adequate. Mentions return type; could specify whether it accepts bundles or single resources, but overall sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage with description 'The FHIR resource to validate'. Description adds no additional parameter meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb 'validate', the resource 'FHIR R6 resource', and the scope 'against structural rules'. Also mentions the return type 'OperationOutcome'. Distinct from sibling tools like fhir_read or fhir_commit_write.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. Usage is implied as a pre-commit check, but no exclusions or contextual hints provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guardrail_conformanceGuardrail Conformance ScorecardA
Read-only
Inspect

Run the guardrail conformance self-test on the connected HealthClaw deployment and return the graded scorecard across seven guardrail properties: PHI redaction, immutable audit, step-up auth, human-in-the-loop, tenant isolation, medical disclaimers, and error fidelity. Uses synthetic data only. Set fresh=true to force a new run instead of the cached result.

ParametersJSON Schema
NameRequiredDescriptionDefault
freshNoForce a fresh probe run instead of the cached (<=10 min old) result

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and destructiveHint false. The description adds that the tool uses synthetic data only and may return cached results (with a 10-minute staleness). It does not mention specific auth requirements or error behavior, but the annotations cover the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first states the main purpose and enumerates the seven guardrail properties; the second explains the lone optional parameter. No wasted words, front-loaded with the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description lists the seven properties in the scorecard, which provides sufficient expectation. It also explains caching and synthetic data. However, it lacks prerequisites or error conditions, which are not critical for a self-test tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single 'fresh' parameter. The description adds context that the cached result is <=10 minutes old and that setting fresh=true forces a new run, which goes beyond the brief schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs a guardrail conformance self-test and returns a graded scorecard across seven specific properties. It distinguishes itself from the numerous sibling tools, none of which perform a similar guardrail check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly indicates when to use (to check guardrail conformance) and includes a note about synthetic data, but does not explicitly state when not to use or discuss alternatives. The uniqueness among siblings reduces the need for explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionnaire_extractExtract Form Data to FHIRA
Destructive
Inspect

SDC $extract — extract FHIR resources from a completed QuestionnaireResponse into a transaction Bundle. Write tier; requires step-up unless dry_run=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoPreview the Bundle without committing
questionnaireNoThe referenced Questionnaire (optional if resolvable by reference)
questionnaire_responseYesCompleted QuestionnaireResponse

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true, and the description reinforces this with 'Write tier' and the dry_run parameter to make it safe. This adds context beyond annotations by clarifying the step-up requirement and the dry_run escape hatch. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences that front-load the purpose and behavioral context. No redundant information; every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description explains the output as a transaction Bundle, which is adequate. It covers input, behavior, and the dry_run option. Slightly more detail on the Bundle structure would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters have descriptions in the schema (100% coverage), so the baseline is 3. The description mentions 'dry_run=true' but does not add significant meaning beyond the schema. The overall context of SDC $extract is helpful but not extra parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool extracts FHIR resources from a completed QuestionnaireResponse into a transaction Bundle, using the SDC $extract operation. This differentiates it from sibling tools like questionnaire_populate, which likely populates rather than extracts. The purpose is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates this is a write-tier operation requiring step-up unless dry_run=true, providing clear context on when to use it and the need for permissions. However, it does not explicitly mention when not to use this tool or suggest alternatives, leaving some implicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

questionnaire_populatePre-fill Health FormA
Read-only
Inspect

SDC $populate — pre-fill a Questionnaire for a subject. Returns a QuestionnaireResponse. Read tier; mints a tenant token for non-public tenants.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionnaireNoInline Questionnaire (overrides questionnaire_id)
questionnaire_idNoStored Questionnaire id
subject_referenceYesSubject reference, e.g. 'Patient/p1'

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=false. Description adds value by stating it mints a tenant token, which is not in annotations, and clarifies the access tier. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with key action and result, no redundant words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, description specifies return type (QuestionnaireResponse). It mentions token minting for non-public tenants, adding context. Could elaborate on when to use inline vs stored questionnaire, but schema handles that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all three parameters with descriptions (100% coverage). Description does not add additional meaning beyond what is in the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'pre-fill a Questionnaire for a subject' with a specific verb and resource, and includes return type and access tier. It distinguishes from sibling 'questionnaire_extract' by focusing on population rather than extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description mentions 'Read tier' indicating safe context, and notes token minting for non-public tenants, but does not explicitly state when not to use or provide alternative tools. Sibling list includes 'questionnaire_extract' which could be an alternative but is not referenced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rx_transfer_requestRequest Prescription TransferA
Read-only
Inspect

Draft a prescription-transfer request: assembles the patient's active medications and stages a phone call to the RECEIVING pharmacy asking it to pull the prescriptions from the current pharmacy (how US transfers actually work). Schedule II medications are refused (never transferable — new prescription required). Returns a draft the patient MUST review; submit with action_commit for the patient's own out-of-band confirmation after they explicitly agree — action_commit does not execute the call itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
medication_namesNoLimit to these medication names (default: all active orders)
to_pharmacy_nameYesReceiving pharmacy name
to_pharmacy_phoneYesReceiving pharmacy phone number
from_pharmacy_nameNoCurrent pharmacy name (optional)
from_pharmacy_phoneNoCurrent pharmacy phone (optional)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and destructiveHint=false, and the description confirms it only creates a draft, not executing any action. It adds behavioral context: Schedule II refusal, need for patient review, and reliance on action_commit for confirmation. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core action. Every sentence adds essential information: purpose, process, constraints, and next steps. No redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description states what is returned (a draft requiring review) and explains the workflow with action_commit. It covers key behavioral aspects (Schedule II refusal) and the open-world hint (patient must review). It could detail the draft format more, but it's complete enough for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds meaning beyond the schema by explaining medication_names as optional limiting and clarifying the roles of from_pharmacy vs to_pharmacy. It also connects parameters to the transfer process, adding value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool drafts a prescription transfer request, explains the US transfer process, and distinguishes from sibling action_commit. It specifies the resource (active medications) and the action (staging a phone call to receiving pharmacy). The purpose is unambiguous and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use action_commit after reviewing the draft, providing a clear workflow. It also warns that Schedule II medications are not transferable, guiding appropriate use. However, it does not explicitly list scenarios where the tool should not be used (e.g., emergency transfers), but the context is sufficient for an AI agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shl_generateGenerate SMART Health LinkAInspect

Generate a SMART Health Link (shlink:/ QR payload) sharing the patient's record with a clinic. Fetches the guardrailed share-bundle from HealthClaw (step-up required — pass _stepUpToken), encrypts it client-side (the SHL server never sees plaintext), uploads ciphertext, and returns the shlink URI, viewer link, and the patient's private manage link. ALWAYS get the patient's explicit consent before generating, and deliver the manage link ONLY to the patient.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoShort label shown in SHL viewers (<=80 chars), e.g. 'Records for Winters Healthcare'. No PHI beyond what the patient approves.
profileNointake = identified record for clinic check-in (default); deidentified = strips name/contact/institutional IDs
patient_idNoOptional patient id filter for multi-patient tenants
expires_in_daysNoLink lifetime in days (default 7, max 90)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial context beyond annotations: step-up token requirement, client-side encryption (SHL server never sees plaintext), ciphertext upload, and returned link types. It also warns about consent and manage link delivery. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: it states the main action, outlines step-by-step, and ends with important usage warnings. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains return values (shlink URI, viewer link, manage link). It provides sufficient context for correct usage, including consent and delivery instructions. The tool complexity is moderate and the description fully addresses it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all parameters are described in the schema. The description does not add additional meaning beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates a SMART Health Link and explains the process: fetching a share-bundle, client-side encryption, upload, and returning URIs. It uses specific verbs and resource, distinguishing it from siblings like FHIR tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit consent requirement and instruction to deliver the manage link only to the patient. It implicitly guides when to use (sharing patient record with a clinic) but lacks explicit alternatives or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sources_checkCheck Data SourcesA
Read-only
Inspect

Survey ALL connected health data sources (Fasten, HealthEx, Health Bank One, MEDENT, Flexpa, Epic/Health Skillz, wearables) at once — returns each source's connection status and the patient's record counts by type. Use when the patient asks what's connected or to check for data across services.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond the readOnlyHint and destructiveHint annotations by explaining that the tool surveys all sources at once and returns per-source status and counts. This behavioral context is not deducible from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: first describes functionality, second gives usage guidance. No filler, every word adds value. Front-loaded with the key action ('Survey ALL').

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter input and no output schema, the description sufficiently explains what the tool does and when to use it. It could be slightly more detailed about the output format, but it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the schema covers all. The description doesn't need to add parameter info, and a baseline of 4 is appropriate as per guidelines.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Survey' and clearly states it checks ALL connected health data sources at once, returning connection status and record counts. This distinguishes it from sibling tools like wearables_sync_status which focus on a single source.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides use cases: 'when the patient asks what's connected or to check for data across services.' While it doesn't list exclusions, the context is clear enough for an agent to decide when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wearables_sync_statusWearables Sync StatusA
Read-only
Inspect

List wearable connections (Garmin, Oura, Polar, Suunto, Whoop, Fitbit, Strava, Ultrahuman) for a tenant, with last sync time, observation count, and status. Use this to tell a patient what's connected, when data last arrived, and surface a connection-management UI (via _meta.ui.resourceUri) so they can connect more providers. Data flows into HealthClaw as FHIR Observations with LOINC codes — agents read it via fhir_search like any other Observation.

ParametersJSON Schema
NameRequiredDescriptionDefault
tenant_idNoTenant to inspect. Defaults to the incoming X-Tenant-Id header.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, providing a solid safety profile. The description adds context about data flowing into HealthClaw as FHIR Observations, but does not detail any additional behavioral traits like error handling or pagination. The added value is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured, starting with the core function, followed by use case and data flow. It is informative without being excessively verbose. A minor reduction for including information about HealthClaw that is not essential for immediate tool usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (one optional parameter, no output schema), the description fully covers purpose, usage context, and output format. It also references the UI resource URI and relates to other tools, making it complete for agent decision-making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 100% of parameters (tenant_id with description). The description does not add any additional semantics beyond what the schema already provides for the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool lists wearable connections with details like last sync time, observation count, and status. It lists supported brands and clearly distinguishes itself from sibling tools like fhir_search by stating its specific use for checking sync status, not reading observations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance: 'Use this to tell a patient what's connected, when data last arrived, and surface a connection-management UI.' It does not explicitly list when not to use or alternatives, but the context is sufficient for an agent to decide when to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 29 tool updatesv1.8.0
    • First observedaction_commit
    • First observedaction_propose
    • First observedaction_status
    • First observedcare_gaps
    • First observedcontext_get
    • First observedcuratr_apply_fix
    • First observedcuratr_evaluate
    • First observedfetch
    • First observedfhir_commit_write
    • First observedfhir_compiled_truth
    • First observedfhir_get_token
    • First observedfhir_interpret_labs
    • First observedfhir_lastn
    • First observedfhir_permission_evaluate
    • First observedfhir_propose_write
    • First observedfhir_read
    • First observedfhir_search
    • First observedfhir_seed
    • First observedfhir_stats
    • First observedfhir_subscription_topics
    • First observedfhir_validate
    • First observedguardrail_conformance
    • First observedquestionnaire_extract
    • First observedquestionnaire_populate
    • First observedrx_transfer_request
    • First observedsearch
    • First observedshl_generate
    • First observedsources_check
    • First observedwearables_sync_status

TDQS

A3.6/5.0
Disambiguation4/5

Most tools have distinct purposes, but some overlap exists, e.g., 'search' vs 'fhir_search' and 'action_commit' vs 'fhir_commit_write'. Descriptions help clarify, but the agent might occasionally select the wrong tool.

Naming Consistency4/5

Tool names follow a mostly consistent snake_case pattern with hierarchical prefixes (fhir_, action_, curatr_). There are a few deviations like 'curatr_apply_fix' mixing product name, but overall pattern is clear.

Tool Count2/5

With 29 tools, the server covers an extremely broad scope (FHIR CRUD, data quality, actions, questionnaires, transfers, wearables, guardrail testing), which is too many for a coherent, focused toolset. Typically, servers with this many tools become unwieldy.

Completeness4/5

The toolset covers major healthcare workflows: resource management, data quality, lab interpretation, care gaps, prescription transfer, and more. Minor gaps exist (e.g., no action cancellation, no direct consent management), but the overall surface is comprehensive.

Maintenance

ActivityActive
ResponsivenessResponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/aks129/HealthClawGuardrails'

If you have feedback or need assistance with the MCP directory API, please join our Discord server