Skip to main content
Glama
Hei33enberg

WhiteIntel MCP Server

by Hei33enberg

@whiteintel/mcp-server

They trace names. We trace who really owns them.

Add to Replit

The corporate-ownership & sanctions intelligence layer for AI agents — built for the agentic era. WhiteIntel turns public-registry and offshore-leak data into MCP-native intelligence primitives — entity search, semantic discovery, ownership-path traversal, sanctions screening, offshore-exposure detection, and fully cited dossiers — so any AI agent can investigate a company, trace its ultimate beneficial owner, and flag risk in one conversation. Your agent isn't querying a database — it's conducting an investigation.

Read the Methodology →

npm CI License: MIT MCP Tools Corpus Sources whiteintel.dev


What's live today

One command, any MCP agent:

npx -y @whiteintel/mcp-server

…starts an MCP server with 24 tools that give any AI agent — Claude Desktop, Cursor, Cline, Windsurf, or your own runtime — full corporate-ownership intelligence: search by name or meaning, trace ownership chains to the UBO, screen sanctions across OFAC/EU/UN/UK, detect offshore layering, pull fully cited dossiers with financials and asset layers, and even purchase deeper intelligence through agent-initiated Stripe checkout. Every claim cited to its source, every edge traced to a registry record.

Tool

What it does

Category

search_entities

Search the corpus (companies + people) by name → entity ids

🔍 Discovery

semantic_search

Meaning-based search (BGE-M3 vector ANN) — find entities by profile, not keywords

🔍 Discovery

find_similar

"More like this" — nearest entities to a known id, for peer discovery and clustering

🔍 Discovery

search_companies

Free-text company-name search → registration number

🔍 Discovery

lookup_company

UK company by Companies House number → record + ownership graph

📋 Lookup

lookup_by_identifier

Resolve by strong id — LEI, OFAC/EU/UN/UK sanctions id, UEN, SEC CIK, KRS, GB-COH, SIREN, Brazil RFB CNPJ

📋 Lookup

get_entity

Full record for one entity + its direct relationships

📋 Lookup

resolve

Batch-resolve names or scheme:value ids → canonical entity ids + confidence

📋 Lookup

list_jurisdictions

Coverage map per country — tier, scope and record depth we hold

🗺️ Coverage

list_asset_coverage

Coverage map per asset class — aircraft, vessels, real estate

🗺️ Coverage

get_dossier

Structured, fully-cited dossier: identity, ownership/UBO chain, risk, provenance

📊 Intelligence

trace_ownership_path

Walk ownership upward to the ultimate beneficial owner

📊 Intelligence

graph_neighbourhood

Every edge within N hops of an entity, both directions — hard-capped, says when the view is partial

🕸️ Graph

graph_path

How two entities are connected — bounded, not exhaustive: found: false is not proof of no link

🕸️ Graph

get_sanctions

Sanctions exposure (OFAC/EU/UN/UK) for entity and its resolved cluster siblings

🛡️ Risk

check_offshore_exposure

Flag sanctioned + secrecy-jurisdiction hops in the ownership chain

🛡️ Risk

get_company_details

UK register detail: address, status, SIC, filings, charges, former names

📋 Lookup

get_financials

Filed UK financials YoY (turnover, profit, net assets, cash, employees)

📊 Intelligence

get_pulse

Live corpus activity feed — recent ownership/control changes, sourced

📊 Intelligence

get_pricing

Full price list + machine-readable purchase flow (static, no network call)

💳 Commerce

buy_dossier

Start a one-off dossier purchase via guest Stripe Checkout → checkout_url

💳 Commerce

get_payment_link

Permanent, reusable Stripe payment links — the artefact you hand to a human

💳 Commerce

claim_dossier

Redeem a paid session for a 90-day access token (idempotent)

💳 Commerce

24 callable tools — 4 Discovery + 4 Lookup + 2 Coverage + 4 Intelligence + 2 Graph + 2 Risk + 4 Commerce + 1 Feed + 1 Pricing. All read-only except buy_dossier (opens Stripe — money moves only when a human completes it) and claim_dossier (redeems an already-paid session). Ids flow between tools: search → get_dossier → trace_ownership_path → get_sanctions.

Related MCP server: ENTIA Entity Verification

Quickstart (60 seconds)

Distribution: the package is on npm — npx -y @whiteintel/mcp-server Just Works.

1. Run it. No key needed — works anonymously on the free tier:

npx -y @whiteintel/mcp-server

2a. Claude Desktop / Cursor — add to your MCP config:

{
  "mcpServers": {
    "whiteintel": {
      "command": "npx",
      "args": ["-y", "@whiteintel/mcp-server"],
      "env": { "WHITEINTEL_API_KEY": "wi_…" }
    }
  }
}

2b. Claude Code CLI:

claude mcp add whiteintel -- npx -y @whiteintel/mcp-server

2c. One-click: add WhiteIntel to your editor at whiteintel.dev/developers.

The env block is optional — omit it to use the anonymous free tier. Set WHITEINTEL_API_KEY=wi_… to authenticate as your plan and lift limits.

Try it

You: "Who ultimately owns Revolut? Check sanctions on the whole chain."

Agent: calls search_entities({ query: "Revolut" })trace_ownership_path({ id })get_sanctions({ id }) for each hop → a fully cited ownership chain with sanctions screening at every level. Done.

You: "Find companies similar to Wirecard and check for offshore exposure."

Agent: calls find_similar({ entity_id })check_offshore_exposure({ id }) → flagged secrecy-jurisdiction hops and sanctioned intermediaries across the peer set.

Agents can pay

An agent can buy the paid depth of a dossier end-to-end, no WhiteIntel account needed:

  1. buy_dossier { tier: "standard" | "premium", entity_id } → returns a Stripe checkout_url. Standard (€39) unlocks the full multi-hop UBO chain + financials; Premium (€99) adds aircraft, sanctioned vessels and property; on HIGH-risk or sanctioned subjects it additionally runs a live adverse-media scan (that scan is gated — it does not run on lower-risk entities).

  2. A human completes payment at the checkout_url — Stripe collects an email and redirects back.

  3. claim_dossier { session_id }{ token, entity_id, tier }. Idempotent; returns 402 until paid.

  4. get_dossier { id, token } → the unlocked, fully-cited dossier JSON. Tokens valid 90 days.

No human at the keyboard right now? Step 1 is the wrong tool: a checkout_url is single-use and expires in 24 hours, so it is dead by the time someone reads your report. Call get_payment_link instead — it returns permanent Stripe links you can paste into a document, a ticket or a message, and append ?client_reference_id=<entity uuid> to bind one to a specific company. Measured 2026-08-11: those links cover the Standard tier only (single / 5 / 25); Premium still goes through buy_dossier.

Check get_pricing first — it returns the full price list plus this flow in machine-readable form.

The corpus

~171.2M entities across 41 fused registries — every claim cited, every edge traced.

Measured 2026-08-23 from whiteintel.dev/api/public/stats (entities = 171,207,760, itself a planner estimate). That endpoint rebuilds its source map by counting registries, so it is always the authority — and a new source shows up there without anyone editing this file.

Source

What

Coverage

OpenOwnership

UK PSCs (Persons with Significant Control)

🇬🇧 Full

GLEIF

Global LEI registry + parent/child ownership relations — nightly refresh scheduled

🌍 Global

ACRA Singapore

Singapore company registry

🇸🇬 Full

ICIJ Offshore Leaks

Panama Papers, Paradise Papers, Pandora Papers

🌍 Offshore

SEC EDGAR

US securities filings + beneficial ownership

🇺🇸 Full

UK Companies House

Full UK register — bulk + live filing stream

🇬🇧 Full

FAA

US aircraft registry (tail numbers → owners)

🇺🇸 Full

France SIRENE

French company register

🇫🇷 Full

Brazil RFB

Brazilian federal revenue — CNPJ register

🇧🇷 Full

Cyprus DRCOR

Cypriot register — officers only (see scope note below)

🇨🇾 Loading

OFAC / EU / UN / UK

Consolidated sanctions lists

🌍 Live

+ 26 more

registries, sanctions lists & UBO registers

🌍 Growing

Cyprus — what it is, and what it is not

Cyprus went to production on 2026-08-11 and is still loading — so we quote no frozen row count here; ask /api/public/stats for the current figure.

Read this before you sell it as Cyprus ownership coverage — it is not. The Cypriot open data release covers the nominal layer only: directors, secretaries and trade-name owners. It contains no shareholders and no beneficial owners. Measured on a sample of the loaded edges, roughly 93% are Directorship (Director, Secretary, Authorised Person, general partner) and the remaining ~7% carry the Ownership schema with role Owner — those are trade-name proprietorships, a sole trader registered behind a business name, not shareholding in a company. An earlier version of this paragraph said there was "not one ownership edge" in the Cyprus data; that was wrong, and it is corrected here rather than quietly deleted, because a claim about what a source does not contain is exactly the kind of sentence a buyer relies on.

The practical consequence is unchanged and is the part that matters: a Cypriot company will typically answer trace_ownership_path and check_offshore_exposure with no_ownership_data. That verdict means we hold no ownership edges for this subject, not this company is cleanly owned. Do not read the 7% as shareholder coverage — it is not.

Cypriot records carry a cy-reg: identifier. lookup_by_identifier does not accept that scheme — reach them with search_entities using juris: "cy".

Contains information from the Cyprus Department of Registrar of Companies and Intellectual Property, licensed under CC BY 4.0.

Semantic search (semantic_search / find_similar) runs over resolved dossier cards using BGE-M3 embeddings; coverage grows as the embedding backfill completes. Measured 2026-08-11 from the endpoint's own coverage payload: 990,055 of a 47,486,969 universe embedded (2.1%), and that slice is ~99.6% risk-listed and ~97% natural persons — so today these two tools behave much more like a sanctions/PEP search than a corpus search, and an empty result usually means "not embedded yet". Lexical search_entities always covers the full corpus; pair it with either of them before drawing a conclusion.

Why WhiteIntel

What's in a name: White + Intel — white as in transparent, open, cited; intel as in intelligence, not data. We don't sell raw records — we sell resolution, traversal, and cited delivery.

Existing corporate-ownership tools were built for compliance analysts clicking web forms. WhiteIntel is the intelligence layer for the agentic era — where the investigator might be a person, an autonomous agent, or an AI workflow, and they all need the same cited, traversed, risk-scored intelligence.

  • Cited, not claimed. Every ownership edge, every sanctions flag, every risk signal is traced to a public-registry record with a real effective date. We don't invent or infer — if a source doesn't say it, we don't.

  • MCP-native, not another API wrapper. Semantic intelligence primitives — not REST endpoints shoe-horned into tool definitions. One command, any MCP agent.

  • Freemium by design. The public corpus is free to explore — no sign-in, no API key, no paywall on search. You pay only for depth: full UBO chains, asset layers, monitoring, and export.

  • Agents can pay. The only MCP server where an agent can investigate a company, decide it needs the paid dossier, buy it via Stripe Checkout, and receive the cited intelligence — end-to-end, no human portal needed.

  • Honest about gaps. An absent edge means "not yet observed", not "does not exist". Investigative decision-support, not a legal determination of beneficial ownership.

  • No lock-in. MIT license. Your agent, your data, your investigation.

Data & honesty

  • Live corpus: ~171.2M entities across 41 fused registries (measured 2026-08-23). Live counts, always authoritative over this file: whiteintel.dev/api/public/stats.

  • Sources are not uniformly deep. A registry in the list means we hold what that registry publishes — which for some jurisdictions is the officer layer, not ownership. Cyprus is the clearest case (see the scope note above). Never read presence in the source table as ownership coverage.

  • An absent edge means "not yet observed", not "does not exist".

  • Investigative decision-support, not a legal determination of beneficial ownership.

  • Semantic search coverage grows as the embedding backfill completes — lexical search always covers the full corpus.

Configuration

Env var

Default

Purpose

WHITEINTEL_API_KEY

(none)

Optional wi_ key (whiteintel.dev → Settings → API keys). Authenticates as your plan, lifts free-tier limits.

WHITEINTEL_API_BASE

https://whiteintel.dev

API origin (SSRF-guarded to whiteintel.dev hosts).

WHITEINTEL_TIMEOUT_MS

30000

Per-request timeout.

Ecosystem

WhiteIntel is part of a growing intelligence platform:

Who's behind this

WhiteIntel is built and directed by @Hei33enberg — a self-funded, independent intelligence project. No venture capital, no data brokers, no compromises on citation integrity.

Swiss governance · Honest by construction

Get on the graph

npx -y @whiteintel/mcp-server     # 24 tools, any MCP agent
  • Install — drop the server into Claude Desktop, Cursor, Cline, Windsurf, or your own runtime (see Quickstart).

  • No key needed — works on the anonymous free tier out of the box.

  • Go deeper — set WHITEINTEL_API_KEY for your plan's full depth.

  • Explore the corpuswhiteintel.dev — free to search, no sign-in.

  • Own itstar the repo, build on the API, or integrate into your agent pipeline. MIT, no lock-in.

Contributing

Issues, PRs, and tool ideas welcome. Start with the CHANGELOG for what's shipped and what's next. If you're building an agent that uses corporate intelligence, we want to hear about it — intel@whiteintel.dev.

Community: GitHub Issues for bugs and features, GitHub Discussions for design and help.

Web: whiteintel.dev · npm: @whiteintel/mcp-server · Releases: GitHub

Privacy Policy

https://whiteintel.dev/privacy

The WhiteIntel MCP server runs locally and calls only https://whiteintel.dev (SSRF-guarded). It sends the query terms you pass to a tool and, if set, your WHITEINTEL_API_KEY. It does not read your files, your conversation history, or your environment beyond WHITEINTEL_API_KEY and WHITEINTEL_API_BASE. Query logs are retained for abuse prevention and are not sold or shared with third parties. Contact: hello@whiteintel.dev

License

MIT © whiteintel.dev

Available Tools

21 tools
buy_dossierAInspect

Start a one-off dossier purchase via guest Stripe Checkout — no WhiteIntel account needed (Stripe collects an email for delivery). Pick a tier ('standard' €39: full UBO chain + financial history · 'premium' €99: additionally itemised assets — vessels, aircraft, securities, real estate) and optionally a bulk pack ('5' or '25' report credits; standard 5×€159 / 25×€599, premium 5×€399 — no premium 25-pack) plus the entity_id (from search_entities) the report is for. Returns checkout_url + next_steps: open the URL so payment can be completed, then feed the session_id from the post-payment redirect to claim_dossier for the access token. See get_pricing for the full price list. WRONG TOOL IF NOBODY IS THERE TO PAY: the session it mints is single-use and expires in 24 hours, so putting this URL in a report or a message read tomorrow hands over a dead link. Use get_payment_link for a permanent, reusable one (standard tier only — Premium is available solely through this tool). And do not fetch checkout_url yourself; it is a card form, so it must be handed to a human.

ParametersJSON Schema
NameRequiredDescriptionDefault
packNoOptional bulk pack (default single). standard: 5=€159 / 25=€599 · premium: 5=€399 (no 25-pack).
tierYesDossier tier: standard (€39) or premium (€99, adds itemised assets).
entity_idNoOptional entity id (from search_entities) the dossier should unlock.
entity_nameNoOptional entity display name, recorded on the Stripe session as an audit trace only — it is NOT displayed anywhere. Since 2026-08-09 the name shown on the invoice and in the delivery email is read from WhiteIntel's own record for entity_id (caller-supplied text is never rendered in mail we send), and the Checkout page shows the Stripe product name. Safe to omit.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses critical behaviors: guest checkout, Stripe collects email, session is single-use and expires in 24 hours, and caller-supplied entity_name is not displayed (safety about audit data). It also warns against fetching the URL. This is comprehensive and goes beyond basics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph but efficiently packed with essential information. Each sentence adds value: pricing, required parameters, return value, error conditions, alternatives, and security warnings. Despite length, it is front-loaded with the core purpose and ends with critical usage caveats. No redundant fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though there is no output schema, the description explicitly states the return (checkout_url + next_steps) and the follow-up workflow (feed session_id to claim_dossier). With full parameter documentation and clear behavior, the description is complete for the tool's complexity. It addresses all necessary aspects for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (100%), so baseline is 3, but description adds value: it explains the pricing tiers in prose, clarifies that entity_id comes from search_entities, and details entity_name's audit-only role and its recent behavior change. It supplements the schema with context that helps correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource and clear purpose: 'Start a one-off dossier purchase via guest Stripe доListening surtout mailing类专业 unsaturated бы粱作为一名大笑mesytetiwxx: product name visible anywhere, product name without, "buy_dossier" not mention in "Buy" or , and line: "Start a one-off dossier purchase..." This differentiates from siblings like get_payment_link (permanent link) and claim_dossier (token retrieval). It clearly states the action and the return of checkout_url, and distinguishes from get_pricing and get_payment_link.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool vs alternatives: 'WRONG TOOL IF NOBODY IS THERE TO PAY...' and 'Use get_payment_link for a permanent, reusable one...'. It also warns not to fetch checkout_url programmatically. This provides clear when/when-not guidance and names alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_offshore_exposureAInspect

Walk the ownership chain upward from an entity and flag, hop by hop, whether each node is sanctioned and/or sits in a secrecy jurisdiction (classic tax-haven / offshore-secrecy country). Returns the chain, the flagged hops, and a structured 4-state verdict — BRANCH ON verdict, NOT on exposed. States: no_ownership_data (we hold zero ownership edges from this entity — NOT a clean verdict, exposure cannot be evaluated), flagged (a sanctioned or secrecy-jurisdiction hit sits on the walked chain), checked_to_max_depth_truncated (walk reached the depth cap with more chain above — a flagged owner may still sit higher, NOT clean), checked_full_clean (the walk ran out of chain before the cap, no flag). Also returns depth_walked (how deep the walk actually reached) and depth_capped. READ depth_capped EVEN WHEN THE VERDICT IS checked_full_clean, because the two co-occur. Measured 2026-08-11 anonymously with max_depth=6: verdict: 'checked_full_clean', depth_walked: 1, depth_capped: true, plan: 'free'. depth_capped: true means A CAP WAS IN FORCE, not that the cap necessarily bit — here the chain genuinely ended after one hop, below the free plan's 2-hop ceiling. The honest report of that response is 'clean over the one hop of ownership we hold, on a walk a free key limits to two', which is what the payload's own note says in prose. Never promote checked_full_clean to 'no offshore exposure' without quoting depth_walked. Anonymous callers walk at most 2 hops however high you set max_depth. Legacy exposed boolean is retained but is only meaningful when verdict='flagged'. Get the id from search_entities or lookup_by_identifier.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id to assess.
max_depthNoMax ownership hops to request (default 6). Anonymous callers are capped at 2 — read `depth_walked` in the response.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral disclosure. It details depth caps, the meaning of depth_capped, the legacy exposed field's limited usefulness, and a concrete anonymous-caller example, giving the agent excellent expectations for edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, but the complexity of the verdict semantics justifies much of the length. It is front-loaded with the core purpose and then walks through states and caveats; however, the measured example and repeated warnings could be tightened without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description thoroughly explains return values: the chain, flagged hops, verdict states, depth_walked, depth_capped, and legacy exposed. It also covers cap behavior and id sourcing, making the tool usable without additional external documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though the schema covers both parameters, the description adds meaningful semantic value: it tells the caller to obtain id via search_entities or lookup_by_identifier, and explains that anonymous callers are capped at 2 hops regardless of max_depth. This goes beyond the schema's basic field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Walk the ownership chain upward from an entity and flag...', which clearly identifies what the tool does. It also distinguishes this tool from siblings like trace_ownership_path or get_sanctions by emphasizing the offshore/sanctions-secrecy verdict semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong interpretive guidance, such as 'BRANCH ON verdict, NOT on exposed', and warns against promoting checked_full_clean to 'no offshore exposure' without quoting depth_walked. It does not explicitly name alternatives or exclusions, but the context for appropriate use is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claim_dossierAInspect

Redeem a paid Stripe Checkout session for a dossier access token. Pass the session_id (cs_…) from the post-payment redirect after buy_dossier. Returns { token, entity_id, tier } — pass the token to get_dossier as its token input for the unlocked report (standard: full UBO chain + financial history · premium: additionally itemised assets). Idempotent: claiming the same session again returns the same grant, so it is safe to retry. Fails with 402 not_paid until the payment has actually completed — wait for the human to finish Checkout, then call again.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesStripe Checkout session id (cs_…) from the success redirect.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description fully carries the burden of behavioral disclosure. It reveals idempotency (same session returns same grant, safe to retry), handles failure (402 not_paid until payment completed), and explains the return structure and tier variations. This is comprehensive and directly addresses operational behaviors an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured paragraph. It front-loads the core purpose, then provides usage, return details, and error handling without unnecessary verbosity. Every sentence adds value, no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description thoroughly explains return values ({ token, entity_id, tier }) and how to use them. It also covers idempotency, retry safety, and error conditions. With no annotations, this is as complete as one could expect for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage for session_id with a clear description. The description adds contextual meaning by explaining where to get the session_id (from the post-payment redirect after buy_dossier) and reiterating the format (cs_...). This goes slightly beyond the schema, so it earns a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: redeem a paid Stripe Checkout session for a dossier access token. It uses a specific verb+resource (claim dossier) and distinguishes itself from sibling tools by referencing buy_dossier and get_dossier, which is exactly the kind of differentiation needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: after the post-payment redirect following buy_dossier. It also tells the agent to pass the returned token to get_dossier, and explains when not to call (before payment completion), including the 402 not_paid failure and the retry guidance. This is explicit usage context with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_similarAInspect

Entities most similar to a given one — the nearest corpus dossier cards ('more like this'), for peer discovery and clustering around a known entity. Pass an entity_id from search_entities. Returns { id, count, hits }, each hit with entity_id, caption, kind, jurisdiction, risk and a similarity score. COVERAGE IS PARTIAL AND SKEWED — it draws on the same embedded slice as semantic_search: 990,055 of a 47,486,969 universe (2.1%), ~99.6% risk-listed and ~97% natural persons, measured 2026-08-11 from the sibling endpoint's own coverage payload. An entity outside that slice returns count: 0 with an empty hits array and HTTP 200 — that is 'not embedded', NOT 'no peers exist', and it is the common case for ordinary companies (verified: BARCLAYS BANK PLC returns zero). Never report an empty result as a finding about the entity. Fall back to semantic_search or search_entities.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoMax hits (default 10).
entity_idYesEntity uuid from search_entities.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and excels: discloses partial coverage (990,055/47,486,969 = 2.1%), skew (99.6% risk-listed, 97% natural persons), behavioral nuance (HTTP 200 with count:0 for non-embedded entities), and a verified example (BARCLAYS BANK PLC). Even warns against misinterpreting empty results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place. The description front-loads the core purpose, then returns format, then critical limitation warnings. It is longer than average but all content is necessary behavioral caveats, and it remains well-structured and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description compensates by detailing the return shape ({ id, count, hits } with per-hit fields). It covers the tool's purpose, limitations, fallback alternatives, and edge-case behavior, making it fully self-contained for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description reinforces that entity_id must come from search_entities, but adds little beyond the schema's existing descriptions for entity_id and k. It does not explain k's effect beyond schema, but no compensation needed due to high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: 'Entities most similar to a given one — the nearest corpus dossier cards ('more like this')'. It explicitly names the intended use case (peer discovery, clustering) and distinguishes itself from semantic_search and search_entities via fallback guidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use: 'for peer discovery and clustering around a known entity' and clear exclusion: coverage is partial/skewed, empty results mean 'not embedded' not 'no peers exist'. Directly names alternatives: 'Fall back to semantic_search or search_entities.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_company_detailsAInspect

Companies House register detail for a UK company by entity id: registered address, status, company type, incorporation date, SIC industry codes, and the filing/compliance layer — accounts type, last-filed and next-due dates (flagged when OVERDUE), confirmation-statement status, outstanding mortgage charges, and former ('also known as') names. Use this for 'where is X registered / what does it file / is it overdue / what was it called before'. Returns { entity, company_details, provenance, note, source } — this is the best-populated of the UK detail tools, measured 2026-08-11 at 45 of 48 sampled UK company entities carrying a non-empty company_details (contrast get_financials at 11 of the same 48). Get the id from search_entities or lookup_by_identifier.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id (a UK company).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return structure ('Returns `{ entity, company_details, provenance, note, source }`'), notes data completeness with a specific measurement date and sample size, and flags overdue statuses. It does not mention potential errors, rate limits, or auth requirements, but for a read-only lookup tool, the provided behavioral context is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but well-structured, front-loading the core purpose and data fields, then usage guidance, return format, and data-quality metric. It is slightly long but every sentence adds value, including the comparative metric against get_financials. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (single parameter, no output schema), the description is quite complete: it lists the data fields, return envelope, usage examples, and data-quality context. It could mention pagination or error behavior, but for a single-id lookup with no output schema, the description covers the essential context well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage for the single parameter 'id' with a description ('Entity id (a UK company).'). The description adds that the id is for a UK company and how to obtain it, but does not add significant new semantics beyond the schema. Baseline 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves Companies House register details for a UK company by entity id, listing specific data fields (registered address, status, company type, incorporation date, SIC codes, filing/compliance layer). It distinguishes itself from siblings by noting it is the 'best-populated of the UK detail tools' and contrasts with get_financials, making its purpose and differentiation explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use this for 'where is X registered / what does it file / is it overdue / what was it called before''. It also tells the agent how to obtain the required id ('Get the id from search_entities or lookup_by_identifier'), and implicitly contrasts with get_financials for financial data, giving clear when-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dossierAInspect

Build a structured, fully-cited intelligence dossier for one entity by id: identity with cross-source linked records (the same real-world entity resolved across ICIJ leaks, GLEIF, registries), ownership/control (direct owners, holdings, and the UBO chain), risk signals, and provenance on every layer. Every claim traces to a source URL. Use this for 'tell me everything about X'. Get the id from search_entities. Free tier shows the first ownership hop + latest financials; pass a one-off purchase token (from claim_dossier, see get_pricing / buy_dossier) or set WHITEINTEL_API_KEY to unlock the full depth.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id (from search_entities).
tokenNoOptional one-off dossier access token (from claim_dossier or the delivery email). A standard token unlocks the full UBO chain + financial history for this entity; a premium token additionally unlocks itemised assets.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It does so by disclosing the free tier limitation ('first ownership hop + latest financials'), the token/API key unlock mechanism, and provenance guarantees ('Every claim traces to a source URL'). It does not cover rate limits or error handling, but provides strong contextual behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a solid paragraph without fluff. Each sentence contributes value: purpose, provenance, use case, id source, and tier/token details. It could be slightly shortened, but it is appropriately sized for a complex paid tool with no annotations or output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex dossier tool with no annotations and no output schema, the description is remarkably complete. It explains what the dossier includes (cross-source records, ownership/control, UBO chain, risk signals), what the user gets on free tier, and how to unlock full depth. It doesn't describe the exact return format, but 'structured dossier' implies sufficient structure for the agent to proceed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds context about where the id comes from (search_entities) and where the token comes from (claim_dossier), plus the environment variable alternative. However, the schema already fully documents both parameters, and the description adds little beyond acquisition channels.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Build a structured, fully-cited intelligence dossier for one entity by id'. It clearly distinguishes this from siblings like search_entities ('Get the id from search_entities') and get_entity by emphasizing the comprehensive dossier nature and the 'tell me everything about X' use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit use case ('Use this for "tell me everything about X"') and a prerequisite ('Get the id from search_entities'). It also explains free vs paid tiers and token acquisition, but does not explicitly name alternative tools or say when NOT to use it, which is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_entityAInspect

Full record for one entity by id: type (company/person), identifiers, jurisdiction, risk level, summary and its direct relationships with provenance. Get the id from search_entities or lookup_company.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It discloses the main output shape, listing supported fields, and clarifies the entity types. However, it does not mention caveats, limits on relation depth, formatting, authentication needs, or what happens if the id does not exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and pack is highly relevant. Every phrase adds something meaningful: the action, the input, the output fields, and the id-source relationship.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description is reasonably complete: it states input, output fields, entity types, provenance, and how to get the id. It could be improved by briefly mentioning what is not included or why this is distinct from get_dossier/get_company_details, but the essential usage is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only documents 'Entity id' but the description adds useful semantic guidance: the id refers to a specific entity full record, and the id is typically derived from search_entities or lookup_company. This helps agents understand what value to pass and how to obtain it without duplicating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly names the tool's action ('Full record for one entity by id') and enumerates the returned fields (type, identifiers, jurisdiction, risk level, summary, direct relationships with provenance). It distinguishes itself from search tools by referring to them as id-sources, establishing a clear get-by-id purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides direct workflow guidance: 'Get the id from search_entities or lookup_company.' This clearly implies the primary use case (fetch a full entity record once id is known), but it does not list explicit alternatives or when-not-to-use scenarios relative to other sibling tools like get_dossier or get_company_details.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_financialsAInspect

Filed financial figures for a UK company by entity id, year-over-year, from Companies House iXBRL accounts: turnover, profit/(loss), net assets, cash, shareholder funds, fixed/current assets, and employee count per reporting period. Use this for 'what are X's revenue / profit / net assets / how many employees'. Returns { entity, financials, note, source }. MOST ENTITIES HAVE NOTHING HERE, AND THAT IS THE NORMAL ANSWER, NOT AN ERROR. Measured 2026-08-11 on a sample of 48 UK company entities drawn from search_entities: only 11 returned any filed period — the other 37 came back HTTP 200 with an empty financials and a note saying so (even BARCLAYS BANK PLC, CH 01026167, has none loaded). Earlier versions of this description called balance-sheet coverage 'broad'; it is not. Within the accounts we DO hold, the per-field skew is real: balance-sheet items and employee counts are the well-populated ones, while turnover and profit are sparse because micro-entities file no profit-and-loss account. Read note before writing 'no revenue' — absent filings and a filed zero are different claims. Get the id from search_entities.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id (a UK company).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full transparency burden and exceeds it: it discloses the source, the exact return shape, that most entities return empty financials with a note (with a concrete measured sample and even a Barclays example), that balance-sheet fields are well-populated while turnover/profit are sparse, and that absent filings must not be conflated with filed zeros. This is unusually rich operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but every sentence earns its place: purpose, use-case, return shape, normal-empty caveat, measured evidence, field-skew warning, and note-reading instruction. It is front-loaded with the core function and flows logically from purpose to usage to interpretation, with no filler or tautology.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description compensates by stating the return structure, enumerating the reported financial fields, explaining the empty-financials behavior, and clarifying the meaning of the note field. For a tool whose main risk is misinterpreting missing data, this is a complete and robust description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes id as 'Entity id (a UK company)' at 100% coverage, so this dimension starts at baseline 3. The description adds meaningful provenance ('Get the id from search_entities') and reinforces that the id must be a UK company entity id, going slightly beyond the schema without needing more detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Filed financial figures for a UK company by entity id' using Companies House iXBRL accounts, and enumerates the exact metrics (turnover, profit/loss, net assets, cash, shareholder funds, fixed/current assets, employee count). This clearly distinguishes get_financials from sibling tools such as get_dossier or get_company_details, which serve broader company-profile purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit 'use this for' clause ('what are X's revenue / profit / net assets / how many employees') and points to search_entities as the source for the id. It does not explicitly name alternative tools for other data types or state a 'when not to use' condition, but the coverage warning strongly informs expectation-setting around use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pricingAInspect

WhiteIntel's price list plus the exact machine flow for buying access. One-off cited dossiers (Standard €39: full UBO chain + financial history · Premium €99: additionally itemised assets), bulk packs (5× / 25× at a discount), subscriptions (Investigator €149/seat·mo, Business €1,900/mo) and the metered API. Returns how_an_agent_buys — buy_dossier opens a Stripe Checkout, a human (or payment-capable agent) pays, claim_dossier mints the access token, and get_dossier with that token returns the unlocked report. Step 0 of that list covers the case with no human present: get_payment_link returns permanent Stripe links you can hand over instead. Static data, no network call — check it before recommending a purchase.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and delivers: 'Static data, no network call' explicitly declares the behavioral contract, alerting the agent that no side effects or costs are involved. It also discloses return semantics ('Returns how_an_agent_buys') and reveals the purchase-side behavior requiring external payment ('buy_dossier opens a Stripe Checkout, a human... pays, claim_dossier mints the access token').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Though long, the description front-loads the key purpose first, then organizes information in digestible parentheticals and enumerations without redundancy. Every element — tiers, discounts, flow, edge case, and behavioral flag — earns its place; the final 'Static data, no network call — check it before recommending a purchase' is a dense, purposeful closer with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 params, no annotations, no output schema), the description is remarkably complete: it covers all pricing products, describes the full order-of-operations among siblings, handles the human-less edge case, and flags its read-only network-free nature. There are no gaps an agent would need clarified to use this appropriately when recommending a purchase.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters with 100% trivial schema coverage, so per the rubric baseline is 4. The description enhances this by specifying what the return payload contains (pricing tiers per product line, discounts, and the purchase flow walkthrough), going beyond the bare schema by describing the data an agent would consume. No parameter documentation burden exists here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause, 'WhiteIntel's price list plus the exact machine flow for buying access,' uses a specific verb-plus-resource phrasing that unmistakably identifies what the tool returns. It also distinguishes itself from siblings by clarifying this is static reference data explaining the purchase flow rather than the buying action itself — the description explicitly differentiates from buy_dossier, claim_dossier, get_dossier, and get_payment_link.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance with 'check it before recommending a purchase' and walks through the full purchase pipeline, including an alternative with a conditional discriminator: 'Step 0 of that list covers the case with no human present: get_payment_link returns permanent Stripe links you can hand over instead.' It even notes the metered API tier and the no-human-present edge case, giving the agent a decision map of when each sibling applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pulseAInspect

The WhiteIntel Pulse activity feed: recent corpus events — ownership/control changes, filed accounts, watchlist and sanctions designations — newest first, each with a source registry. Use this to answer 'what changed recently' or to monitor the corpus. All four kinds are live and all four are cited: measured 2026-08-11, 100 of 100 rows carried a source_url for every kind. THE UNFILTERED FEED IS NOT A BALANCED MIX. It applies no kind filter, but it is ordered by ingest recency, so whichever loader ran last fills the head of it. Measured 2026-08-11: the newest 100 rows of the default feed were 100% kind: 'ownership', all from one registry. If you want a particular stream — or any sense of proportion between them — pass kind and do not infer 'nothing else happened' from the default page.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOptional: filter by event kind. Registries measured on the newest 100 rows of each stream, 2026-08-11 — they move as loaders run, so read each row's `source_registry` rather than trusting this note. `ownership` = registry-recorded control changes; today 100/100 came from `borme` (Spain's Boletín Oficial del Registro Mercantil), NOT GLEIF as earlier versions of this description claimed. `filing` = UK Companies House accounts (100/100 `companies_house`). `watchlist` = OpenSanctions non-sanctions listings — politically-exposed persons, criminal/wanted entries and procurement debarments, not PEPs alone (100/100 `opensanctions`). `sanction` = a designation landing on a sanctions list, e.g. OFAC SDN (100/100 `opensanctions`).
limitNoMax events (default 40).
sinceNoOptional sync cursor (ISO-8601): pass the next_since from your last response to get only events ingested after it — poll this to monitor what's new.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to lean on, the description carries the full burden and delivers extraordinary disclosure: the feed is 'NOT A BALANCED MIX', it is 'ordered by ingest recency', and it reveals a measurement that the newest 100 rows were '100% kind: ownership, all from one registry' (2026-08-11). It even corrects prior misinformation ('NOT GLEIF as earlier versions of this description claimed') — the kind of bias disclosure that would be invisible to an agent without this text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description front-loads the core definition, then usage, then warnings — good structure with strong leading content. However, the measurement statistics (100/100, dates, percentages) appear twice: once in the main prose and again inside the `kind` enum descriptions, creating slight redundancy. Every sentence otherwise earns its place; the uppercase emphasis is effective but the near-duplicate measurement reports could be consolidated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description fully compensates: it covers what is returned (events with a source registry), ordering semantics, parameter behavior, per-kind meaning, and dangerous default-bias caveats. For a monitoring/feed tool of moderate complexity, there is nothing essential an agent needs to know that the description omits — it even notes when defaults 'move as loaders run,' setting correct expectations about non-determinism.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description goes far beyond, defining each enum value with measured composition — e.g., 'watchlist = OpenSanctions non-sanctions listings — politically-exposed persons, criminal/wanted entries and procurement debarments, not PEPs alone' — and clarifies the registry source (borme/companies_house/opensanctions). `since` is meaningfully framed as a sync cursor for polling rather than a plain date filter. This is exactly the value the schema's one-line hints forfeit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a crisp noun phrase identifying the exact resource and behavior: 'The WhiteIntel Pulse activity feed: recent corpus events... newest first, each with a source registry.' It enumerates the four event kinds, states the ordering, and notes provenance — a specific verb+resource that is impossible to confuse with the sibling search/lookup/graph tools despite no sibling being named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-to-use: 'Use this to answer "what changed recently" or to monitor the corpus.' It also supplies when-not-to-infer guidance ('do not infer "nothing else happened" from the default page') and instructs the agent to pass `kind` when it wants a particular stream. This directly helps the agent choose between this and the search/entity siblings, even without naming them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sanctionsAInspect

Return an entity's screening exposure for the entity AND its resolved cluster siblings, each with a source URL. IT IS NOT SANCTIONS-ONLY, DESPITE THE NAME — read each row's signal_type. Measured 2026-08-11: BARCLAYS BANK PLC came back sanctioned: false with one signal of signal_type: 'crime' (severity HIGH, source_list opensanctions_crime, a criminal/wanted listing reaching it via its cluster). Only signal_type: 'sanctioned' rows are sanctions designations, and only those reliably carry list and regime — on the crime row both were null, so do not read a null list as missing data. Two consequences: a sanctioned: false response can still contain a HIGH-severity adverse finding you must report, and 'no sanctions signal' (what the top-level flag and note describe) is not 'nothing found'. Response splits the top-level flag: sanctioned_self = a direct listing ON this entity; sanctioned_via_cluster = the flag reaches it ONLY via a cross-source cluster sibling (~2.3% false-positive tail on UK OpenOwnership resolution — treat cluster-only hits as a lead until you verify the sibling really is the same real-world party). The aggregate sanctioned (self OR cluster) is preserved for back-compat. Get the id from search_entities or lookup_by_identifier.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesEntity id.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals cluster expansion, signal_type semantics, null list/regime behavior, the distinction between sanctioned_self and sanctioned_via_cluster, the false-positive tail, and back-compat behavior. This is far beyond what annotations would have provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: it front-loads the core function, then delivers critical caveats, a concrete measured example, flag semantics, and id-source guidance. The structure is logical and dense without redundancy, which is appropriate given the tool's misleading name and nuanced behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the absence of an output schema, and the absence of annotations, the description is remarkably complete. It covers return contents, signal types, null handling, flag meanings, false-positive risk, and how to obtain the required id. An agent has enough information to invoke the tool and interpret its results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents the single `id` parameter with 100% coverage. The description adds useful semantic value by specifying that the id should come from search_entities or lookup_by_identifier, and by framing the id as an entity id. This goes beyond the schema's bare 'Entity id' description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return an entity's screening exposure for the entity AND its resolved cluster siblings, each with a source URL.' It also explicitly distinguishes the tool from its misleading name by clarifying it is not sanctions-only and by directing attention to `signal_type`. This clearly separates it from sibling tools like get_entity or lookup_by_identifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong contextual guidance: it explains what the tool returns, warns that a `sanctioned: false` response can still contain adverse findings, and tells the agent to get the id from search_entities or lookup_by_identifier. It does not explicitly enumerate when not to use the tool versus alternatives, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_neighbourhoodAInspect

Return every ownership/control edge within a bounded number of hops of one entity, in BOTH directions: who it controls, who controls it, and their neighbours. Use it to answer 'what sits around this company?' — the wider view that trace_ownership_path (upward only) does not give. Hard-capped in the database: depth 3, 300 edges, and at most 25 edges followed per entity per direction per hop. READ THE DEPTH FIELDS IN THE RESPONSE — DO NOT ASSUME YOU GOT THE DEPTH YOU ASKED FOR. There is no field called depth any more, and that rename is deliberate: the old depth was the CLAMPED REQUEST, never the depth walked, and it was being read as a promise. The response now carries depth_requested (what your plan allowed), depth_walked (measured off the returned edges' own hop numbers — the only depth that is actually proven), depth_capped, and completeness. Measured 2026-08-11 on an anonymous caller: depth=3 requested returned depth=2 with depth_capped=true, because the free plan caps every walk at 2 hops. Any sentence you write about what is or is not around this entity must be scoped to the RETURNED depth. EDGE COUNTS FELL BY UP TO 2.7x ON 2026-08-11 AND NOTHING WAS LOST — read this before you treat it as the corpus shrinking. Until that date the walk emitted the same edge two and three times at depth 2 or more, edge_count counted the duplicated list, and the duplicates were charged against your edges budget. Measured on identical requests before and after the fix: 72 -> 27, 29 -> 13, and at the maximum budget 300 rows holding 285 real edges -> 300 rows holding 300. So a call you made yesterday and repeat today can return far fewer edges for the same subject: the smaller number is the true one, and your budget now buys real edges. One consequence worth knowing: at depth 1 a root can drop from 4 edges to 2, because the registry genuinely holds rows that are identical in every field this endpoint returns and the response has no way to represent the difference. That is also a correction, not a loss. truncated: true plus a plain-language truncation_note does work and does mean the edge budget ran out (verified with edges=10); that is NORMAL for hub entities (the corpus holds single nodes with more than 22,000 edges) and means the picture is partial, not wrong. Each edge carries origin: 'registry' (observed in a source registry) or 'derived'/'curated'/'asserted' (inferred by WhiteIntel). Get the root id from search_entities or resolve.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootYesRoot entity uuid.
depthNoHops to walk (default 2). A REQUEST, not a guarantee — the plan caps it (anonymous callers measured at 2 hops) and the response's `depth_walked` is the authority — it is measured from the edges that came back, not echoed from your request.
edgesNoEdge budget (default 120). Lower it for a legible picture, raise it for completeness.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full weight. It discloses hard caps, depth clamping, deduplication changes, truncation behavior, edge-count expectations, and a warning not to trust requested depth. This is far more transparent than typical tool descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and sibling distinction, which is helpful. However, it becomes quite long and mixes historical measurements, explicit troubleshooting notes, and warnings in a way that requires careful parsing; a tighter summary would improve scanability without losing key caveats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description provides important response semantics: depth_requested, depth_walked, depth_capped, completeness, truncation, edge_count, and edge origin. It also covers edge cases around hubs and duplicate records, making the tool safe to invoke even under complex conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers 100% of parameters, so the baseline is 3. The description adds meaningful semantic context for the `depth` and `edges` parameters (request vs actual depth, budget and deduplication) and points to source endpoints for `root`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states explicitly that it returns ownership/control edges within a bounded number of hops in both directions, with entity neighbours. It also distinguishes itself from the sibling `trace_ownership_path` by noting that tool is upward-only, making the tool's purpose and scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear use case: 'what sits around this company?' and explicitly contrasts the wider neighbourhood view with `trace_ownership_path` (upward only). It also tells the agent how to obtain the root id via `search_entities` or `resolve`, so invocation prerequisites are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_pathAInspect

Find how two entities are connected: a bounded breadth-first search over ownership and control edges in both directions, returning the ordered hops from one to the other. WARNING, AND IT CHANGES HOW YOU MUST REPORT THE RESULT: this search is BOUNDED, NOT EXHAUSTIVE. At most 15 edges are followed per entity, per direction, per hop, so a genuine connection running through a heavily-connected intermediary can be missed. found: false means NO PATH WAS FOUND WITHIN THOSE BOUNDS — it is NOT evidence that the two entities are unconnected, and must never be reported as a clean result. The response always carries exhaustive: false, a structured verdict (e.g. 'connected_within_bounds') and a bounds_note restating this. AND THE DEPTH YOU GET IS NOT THE DEPTH YOU ASK FOR: the response echoes its own max_depth plus depth_capped, and those are the authority. Measured 2026-08-11 anonymously — max_depth=3 and max_depth=4 both came back as max_depth: 2, depth_capped: true, plan: 'free'. So a free-tier found: false is a two-hop negative however many hops you requested; say two hops, not four.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesEnd entity uuid.
fromYesStart entity uuid.
max_depthNoMax hops to request (default 3). Reduced by the plan — anonymous callers measured at 2 — so read the response's `max_depth` and `depth_capped`. Depth 4 is measurably slower on densely connected entities; request it deliberately.

TDQS

A3.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It excels here, detailing the bounded BFS, non-exhaustive nature, depth-capping behavior, free-tier plan limitation, and the meaning of found:false. It also discusses response fields like exhaustive, verdict, bounds_note, and depth_capped, giving the agent comprehensive insight into runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is verbose and heavily repetitive, with multiple ALL-CAPS warnings and run-on sentences. While every sentence adds some information, the structure is not concise and would benefit from tighter organization. The core message could be delivered in half the length without losing impact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool and the lack of annotations or output schema, the description is highly complete. It covers edge cases (bounded search), failure semantics (found:false), response structure (verdict, bounds_note, depth_capped), plan limitations, and performance characteristics, ensuring the agent is well-informed about all important behaviors.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already documents all parameters, so the baseline is 3. The description notably enhances understanding of max_depth by warning that the response's max_depth may be reduced and advising deliberate use of depth 4. However, it adds no new meaning for the from/to parameters beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Find how two entities are connected' via a bounded breadth-first search over ownership and control edges, returning ordered hops. This is a specific verb+resource description, but it does not explicitly distinguish itself from sibling tools like trace_ownership_path or graph_neighbourhood.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

While the description provides extensive operational warnings (e.g., bounded search, not to report found:false as clean), it offers no explicit guidance on when to choose this tool over siblings or what alternatives exist. The usage context is implied but no exclusions or comparisons are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_by_identifierAInspect

Resolve an entity by a strong external identifier instead of a name — a LEI, OFAC SDN uid, EU/UN/UK sanctions id, Singapore UEN, SEC CIK, Polish KRS, UK Companies House number, French SIREN, or Brazil RFB CNPJ. Returns the single resolved entity (id, type, jurisdiction, identifier, risk) so you can pivot into get_entity / get_dossier / get_sanctions. Use this when you already hold a registry id and want the corpus node behind it. All eleven schemes were exercised against production on 2026-08-11 and every one resolved a real entity — no scheme in this enum is decorative. DISTINGUISH THE TWO FAILURE MODES: an unsupported scheme returns HTTP 400 with error: 'bad_request' and the accepted set spelled out in detail, whereas a supported scheme whose value we simply do not hold returns HTTP 404 error: 'not_found'. A 404 is a statement about the corpus, not about the tool — fall back to search_entities. NOT every identifier you may see in a response is resolvable here — the enum below is the complete accepted set and the route hard-rejects anything else with a 400. In particular Cyprus records carry a cy-reg: identifier that this tool does NOT accept (verified: cy-reg → 400), and neither is the cusip: seen on US securities rows: reach Cypriot companies with search_entities using juris='cy'.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesThe identifier value (e.g. an LEI, an OFAC SDN uid, a Companies House number).
schemeYesIdentifier scheme: lei | ofac | eu | un | uk | uen | sec | krs | gb-coh | siren (French SIREN, 9 digits) | br-cnpj (Brazil RFB CNPJ; accepts 8-digit root or full 14-digit form 12.345.678/0001-95).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses return shape (single entity with id, type, jurisdiction, identifier, risk), distinguishes two failure modes (400 for unsupported scheme, 404 for not found), and states that all schemes were production-tested. It also explicitly documents validation behavior (complete enum, hard-reject with 400).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, usage, failure modes, and exclusions are each addressed. It is front-loaded with the core purpose and structured logically, avoiding redundancy despite its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 schemes, two error modes, non-accepted identifiers) and no output schema, the description is remarkably complete. It covers return value shape, error handling, fallback guidance, and test validation date, leaving no major gap for an agent to misinterpret.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining the complete accepted set, giving examples of non-accepted identifiers, and detailing error responses tied to parameters. However, much of the parameter detail (e.g., br-cnpj formats) is already in the schema, so the incremental addition is moderate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it resolves an entity by strong external identifier, lists specific identifier types, and explicitly distinguishes from search_entities as fallback. The verb 'resolve' and resource 'entity' are specific, and it scopes to single entity results, differentiating it from sibling tools like lookup_company which uses name-based lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use ('when you already hold a registry id'), when-not-to-use with fallback ('A 404 is a statement about the corpus, not about the tool — fall back to search_entities'), and names alternatives (search_entities for Cyprus ids). Also excludes unsupported identifiers (cy-reg, cusip) with concrete advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_companyAInspect

Look up a UK company by its Companies House registration number and return the company record plus a ready-built ownership graph (officers, persons of significant control, parent/subsidiary edges). Pass the number verbatim — do not strip leading zeros (e.g. 09446231, SC123456).

ParametersJSON Schema
NameRequiredDescriptionDefault
numberYesUK Companies House registration number.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the disclosure burden. It clearly states the input constraint behavior (verbatim number, preserve leading zeros) and the return behavior (company record plus ownership graph with specified edge types). It does not discuss missing-company behavior, but for a lookup tool this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the main action, then provides the key nuance (leading zeros) and representative examples.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one required parameter, full schema coverage, and no output schema, the description provides a clear picture of both input and output expectations. It names the graph components, making the tool sufficiently complete for selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers the single parameter, so the baseline is 3. The description adds meaningful semantics beyond the schema: the number must be passed verbatim, leading zeros must not be stripped, and concrete examples are given. This helps the agent handle real-world Companies House identifiers correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Look up a UK company by its Companies House registration number.' It also distinguishes itself from siblings by explicitly promising a 'ready-built ownership graph' (officers, PSCs, parent/subsidiary edges), which differentiates it from search_companies or get_company_details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the primary use case obvious: pass an exact Companies House registration number. It gives clear input guidance ('Pass the number verbatim — do not strip leading zeros') with examples. It does not explicitly mention when a sibling like search_companies should be used instead, but the context is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolveAInspect

Batch-resolve a list of company names or strong identifiers (scheme:value — lei, siren, gb-coh, uen, br-cnpj, sec, ofac, eu, un, uk, krs) to canonical WhiteIntel entity ids in ONE call. Each result carries a confidence: 'exact' (identifier match) or 'name' (top name hit); an unmatched row comes back as { match: null, confidence: null }, so check for it rather than assuming positional success. Use this to enrich a whole list — suppliers, counterparties, a portfolio — without one lookup per row. Then feed the ids into get_dossier / trace_ownership_path / get_sanctions. Up to 25 items anonymously (a 26th returns HTTP 400 with the limit spelled out), 100 with WHITEINTEL_API_KEY. TREAT confidence: 'name' AS A CANDIDATE, NOT A RESOLUTION. It is the top lexical hit and nothing more — measured 2026-08-11, the query 'Tesco' resolved to a FRENCH company literally named TESCO (fr-siren:454067281), not Tesco PLC, while 'gb-coh:00445790' resolved 'exact' to TESCO PLC. Confirm a 'name' match's jurisdiction and identifier before you attach it to a real counterparty; pass an identifier whenever you hold one.

ParametersJSON Schema
NameRequiredDescriptionDefault
queriesYesNames or scheme:value identifiers, e.g. ["Tesco", "siren:552081317", "lei:213800...", "gb-coh:00445790"].

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses behavior thoroughly: it explains confidence levels ('exact' vs 'name'), handling of unmatched rows (returns `{ match: null, confidence: null }`), and warns that 'name' matches are candidate hits, illustrated with a concrete Tesco example. This transparency is essential since no annotations are provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is excessively verbose and repetitive. Key information (e.g., the Tesco example, the confidence semantics, and the limits) is stated multiple times, significantly bloating the text. It could be condensed to a few sentences without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides comprehensive context, including usage scenarios, limits, and edge cases. However, the completeness is marred by redundancy; while all necessary details are present, the over-explanation detracts from efficiency. A more concise version would achieve the same completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explains the 'queries' parameter in detail: it accepts a list of company names or strong identifiers in 'scheme:value' format, with examples like 'siren:552081317' and 'gb-coh:00445790'. This fully clarifies the expected input structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: batch-resolving a list of company names or strong identifiers to canonical WhiteIntel entity IDs. It also mentions the output format and confidence levels, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises when to use the tool: 'Use this to enrich a whole list — suppliers, counterparties, a portfolio — without one lookup per row.' It also notes the batch size limits (25 anonymously, 100 with API key), providing clear operational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_companiesAInspect

Free-text company-name search against UK Companies House. Use this to resolve a company NAME into the registration number that lookup_company needs.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesCompany name or fragment.
limitNoMax results (default 8).

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the transparency burden. It discloses that this is a free-text name search scoped to UK Companies House, but it does not describe return format, ordering, pagination, or no-match behavior. It adds some useful context but is not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences. The first gives the core action and scope, the second explains the purpose and relationship to a sibling tool. Every word earns its place, and the key information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter search tool with no output schema, the description is sufficiently complete: it explains the input, the source, and the intended downstream use. It could benefit from mentioning the response shape, but this is not a major gap given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides full descriptions for both parameters (q and limit) with 100% coverage. The description adds no extra parameter-specific semantics, so it stays at the baseline for schema-covered parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('search') and resource ('UK Companies House') and defines the output as resolving a company name into the registration number. It also ties directly to the sibling tool 'lookup_company', making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly tells when to use this tool: to resolve a name to a registration number needed by lookup_company. However, it does not explicitly mention when not to use it or how it differs from search_entities, so it lacks full exclusion/alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_entitiesAInspect

Search every node in the live WhiteIntel corpus — companies AND people — by name, across all fused sources. This is the lexical search and it always covers the FULL corpus, so it is the fallback whenever semantic_search comes back thin. Returns entity ids you then pass to get_entity or trace_ownership_path. Each hit's source says whether it came from the resolved corpus or a live registry passthrough — it does NOT name the originating registry. For that provenance call get_entity, whose entity.registry_profile names the source register when we hold one — measured 2026-08-11 it was populated on 22 of 32 sampled entities, so expect null sometimes and fall back to linked_records[].registry and connections[].source — or get_dossier, which cites per-record source URLs. Use juris to scope to a country (e.g. gb, ky, us, cy). Reach into the non-UK sources is verified, not assumed: a name search for 'PETROLEO BRASILEIRO' returned FR (siren), BR (lei and br-cnpj) and US (cusip) rows in one response, 2026-08-11.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesEntity name or fragment.
riskNoOptional: filter by risk level.
typeNoOptional: filter by entity kind.
jurisNoOptional: filter by jurisdiction code (e.g. gb, ky, us).
limitNoMax results (default 20).

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that the source field indicates resolved vs live passthrough, does not name the originating registry, suggests get_entity for that, and notes the fallback options. Also provides a concrete example with verified non-UK sources, showing transparency about data coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is overly verbose, containing a long digression about provenance, specific measurement dates, and an example. While the initial sentence is clear, the subsequent details could be streamlined to improve conciseness without losing essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and many siblings, the description covers the primary use case, fallback behavior, result handling, and a key parameter. It could mention potential errors or the exact output format, but it is largely complete for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description only adds extra context for the juris parameter (scoping to a country) and does not elaborate on q, risk, type, or limit. Since the schema already provides descriptions for all parameters (coverage 100%), the added value is limited, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it searches all nodes (companies AND people) by name across all fused sources, and distinguishes it from semantic_search by calling it the fallback lexical search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to use it (fallback when semantic_search returns thin results), how to use the returned IDs (pass to get_entity or trace_ownership_path), and mentions the juris parameter for scoping.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_ownership_pathAInspect

Walk the ownership graph upward from a root entity and return the ordered hops connecting it to the ultimate beneficial owner. Use this to answer 'who ultimately controls X?'. Get the root id from search_entities. THE HOP AT THE TOP OF THE LIST IS NOT NECESSARILY THE ULTIMATE OWNER, AND max_depth IS A REQUEST, NOT A PROMISE. Measured 2026-08-11 anonymously: max_depth=6 came back as max_depth: 2, depth_capped: true, plan: 'free' — the walk stopped two hops up and the payload said so only in those two fields. So before you name a UBO, compare hop_count with the RETURNED max_depth and check depth_capped: if the walk was capped and the topmost owner still has owners, you have found an intermediate holder, not the beneficial owner. A paid API key walks deeper. Shape: a single flat hops array (each hop from/fromName/to/toName/role/share/source), not one array per branch. as_observed is a standing caveat: edges carry the date we OBSERVED them in a registry, not a validity period — we hold no ownership end dates, so a link shown here may already have ended.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootYesRoot entity id to trace from.
max_depthNoMax hops to request (default 6). The plan lowers it — anonymous callers measured at 2 — so trust the response's `max_depth` / `depth_capped`, not this value.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description reveals critical behavior: max_depth is 'a REQUEST, not a PROMISE,' top hop is not necessarily UBO, and includes a measured example of depth capping (max_depth=6 returned as max_depth:2, depth_capped:true). It also discloses the standing caveat on `as_observed` dates, covering data-validity limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries unique information: purpose, use case, root source, critical warnings, response shape, and caveats. It is front-loaded with the main action, but the density makes it slightly less concise than shorter equivalents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description provides the full response shape (flat `hops` array with fields from/fromName/to/toName/role/share/source), key response fields (`hop_count`, `max_depth`, `depth_capped`), and interpretation guidance. It covers edge cases and data caveats, making it fully contextual.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers both parameters thoroughly (100% coverage), so baseline is 3. The description adds a concrete measured example and the instruction to fetch root from search_entities, enriching semantics but not fundamentally changing the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with 'Walk the ownership graph upward from a root entity and return the ordered hops' – a specific verb+resource+output. It further clarifies the use case with 'who ultimately controls X?' and distinguishes from general graph tools like graph_path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Use this to answer ''who ultimately controls X?''' and instructs 'Get the root id from search_entities.' It also provides detailed guidance on handling depth capping and verifying UBO, effectively telling the agent when to trust the result.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 21 tool updatesv0.7.7
    • First observedbuy_dossier
    • First observedcheck_offshore_exposure
    • First observedclaim_dossier
    • First observedfind_similar
    • First observedget_company_details
    • First observedget_dossier
    • First observedget_entity
    • First observedget_financials
    • First observedget_payment_link
    • First observedget_pricing
    • First observedget_pulse
    • First observedget_sanctions
    • First observedgraph_neighbourhood
    • First observedgraph_path
    • First observedlookup_by_identifier
    • First observedlookup_company
    • First observedresolve
    • First observedsearch_companies
    • First observedsearch_entities
    • First observedsemantic_search
    • First observedtrace_ownership_path

TDQS

A4.2/5.0
Disambiguation4/5

Most tools have clearly distinct input/output types (e.g., lookup_company vs search_entities vs lookup_by_identifier), and the detailed usage notes remove almost all ambiguity. Minor overlap exists between search_companies and search_entities for UK names, and among graph traversal tools, but descriptions explicitly direct agents to the correct tool.

Naming Consistency4/5

The majority of tools follow a verb_noun pattern (get_entity, trace_ownership_path, buy_dossier), but a few exceptions—graph_neighbourhood, graph_path, semantic_search, find_similar, resolve—break the pattern. These are minor deviations in an otherwise consistent naming scheme.

Tool Count4/5

21 tools is on the higher end, but each tool serves a distinct function across entity lookup, graph analysis, sanctions screening, financials, and purchasing. The payment-related tools (pricing, buy, link, claim) add count but form a logical workflow. The number is appropriate for the broad domain, though slightly heavier than typical.

Completeness5/5

The toolset covers the full range of entity intelligence: multiple search/resolution methods, deep dives (dossier, details, financials), graph analysis (paths, neighborhoods, UBO tracing), sanctions and offshore checks, activity feed, similarity search, and a complete purchase flow. No critical operation appears missing for the stated purpose.

Maintenance

ActivityActive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides real-time company verification and corporate intelligence by accessing global registries like UK Companies House, Singapore ACRA, and OpenCorporates. It enables AI agents to perform KYC tasks, retrieve company profiles, and conduct automated risk assessments for due diligence workflows.
    249
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Unmodified government company data from 27 registries, live. Cross-border UBO chain walker for AI agents. 60+ tools covering GB, IE, NO, FR, DE, NL, PL, BE, CH, LI, MC, IM, IS, CY, AU, NZ, CA, TW, HK, MY, FI, CZ, ES, IT, KR, US — raw upstream fields preserved, no LLM extraction.
    10
    17
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to access official European business data across 15 EU countries, including company lookups, VAT validation, sanctions screening, and KYB reports.
    11,502
    1
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Hei33enberg/WhiteIntel-OS'

If you have feedback or need assistance with the MCP directory API, please join our Discord server