WhiteIntel MCP Server
@whiteintel/mcp-server
Они отслеживают имена. Мы отслеживаем, кому они на самом деле принадлежат.
Слой корпоративной разведки по собственности и санкциям для ИИ-агентов — создан для эпохи агентов. WhiteIntel превращает данные публичных реестров и офшорных утечек в MCP-нативные примитивы интеллекта — поиск сущностей, семантическое обнаружение, обход цепочек владения, санкционный скрининг, выявление офшорной экспозиции и полностью цитируемые досье — так что любой ИИ-агент может расследовать компанию, проследить её конечного бенефициарного владельца и выявить риски в одном диалоге. Ваш агент не запрашивает базу данных — он проводит расследование.
Что доступно уже сегодня
Одна команда, любой MCP-агент:
npx -y @whiteintel/mcp-server…запускает MCP-сервер с 21 инструментом, который даёт любому ИИ-агенту — Claude Desktop, Cursor, Cline, Windsurf или вашему собственному рантайму — полную разведку по корпоративному владению: поиск по имени или смыслу, обход цепочек владения до UBO, санкционный скрининг по OFAC/EU/UN/UK, выявление офшорных слоёв, получение полностью цитируемых досье с финансовыми данными и слоями активов, а также покупку более глубокой аналитики через инициированный агентом Stripe checkout. Каждое утверждение цитируется с указанием источника, каждое ребро прослеживается до записи реестра.
Инструмент | Что делает | Категория |
| Поиск по корпусу (компании + люди) по имени → идентификаторы сущностей | 🔍 Обнаружение |
| Поиск по смыслу (BGE-M3 векторный ANN) — находите сущности по профилю, а не по ключевым словам | 🔍 Обнаружение |
| «Похожие» — ближайшие сущности к известному id, для поиска аналогов и кластеризации | 🔍 Обнаружение |
| Поиск компаний по свободному тексту → регистрационный номер | 🔍 Обнаружение |
| Британская компания по номеру Companies House → запись + граф владения | 📋 Просмотр |
| Разрешение по сильному идентификатору — LEI, OFAC/EU/UN/UK санкционный id, UEN, SEC CIK, KRS, GB-COH, SIREN, Brazil RFB CNPJ | 📋 Просмотр |
| Полная запись одной сущности + её прямые связи | 📋 Просмотр |
| Пакетное разрешение имён или id вида | 📋 Просмотр |
| Структурированное, полностью цитируемое досье: идентичность, цепочка владения/UBO, риск, происхождение | 📊 Аналитика |
| Обход владения вверх до конечного бенефициарного владельца | 📊 Аналитика |
| Все рёбра в пределах N шагов от сущности, в обе стороны — с жёстким лимитом, сообщает, когда вид неполный | 🕸️ Граф |
| Как связаны две сущности — ограниченно, не исчерпывающе: | 🕸️ Граф |
| Санкционная экспозиция (OFAC/EU/UN/UK) для сущности и её разрешённых кластерных соседей | 🛡️ Риск |
| Выявление санкционированных + юрисдикций с секретностью в цепочке владения | 🛡️ Риск |
| Детали британского реестра: адрес, статус, SIC, подачи, обременения, прежние названия | 📋 Просмотр |
| Поданные британские финансовые данные по годам (оборот, прибыль, чистые активы, денежные средства, сотрудники) | 📊 Аналитика |
| Живая лента активности корпуса — недавние изменения владения/контроля, с источниками | 📊 Аналитика |
| Полный прайс-лист + машиночитаемый процесс покупки (статично, без сетевых вызовов) | 💳 Коммерция |
| Начать разовую покупку досье через гостевой Stripe Checkout → | 💳 Коммерция |
| Постоянные, многоразовые Stripe-ссылки для оплаты — артефакт, который вы передаёте человеку | 💳 Коммерция |
| Активировать оплаченную сессию для 90-дневного токена доступа (идемпотентно) | 💳 Коммерция |
21 вызываемый инструмент — 4 Обнаружение + 4 Просмотр + 4 Аналитика + 2 Граф + 2 Риск + 3 Коммерция + 1 Лента + 1 Прайс. Все только для чтения, кроме buy_dossier (открывает Stripe — деньги двигаются только когда человек завершает оплату) и claim_dossier (активирует уже оплаченную сессию). Идентификаторы перетекают между инструментами: search → get_dossier → trace_ownership_path → get_sanctions.
Related MCP server: ENTIA Entity Verification
Быстрый старт (60 секунд)
Распространение: пакет на npm —
npx -y @whiteintel/mcp-serverработает из коробки.
1. Запустите. Ключ не нужен — работает анонимно на бесплатном тарифе:
npx -y @whiteintel/mcp-server2a. Claude Desktop / Cursor — добавьте в конфиг MCP:
{
"mcpServers": {
"whiteintel": {
"command": "npx",
"args": ["-y", "@whiteintel/mcp-server"],
"env": { "WHITEINTEL_API_KEY": "wi_…" }
}
}
}2b. Claude Code CLI:
claude mcp add whiteintel -- npx -y @whiteintel/mcp-server2c. В один клик: добавьте WhiteIntel в свой редактор на whiteintel.dev/developers.
Блок env необязателен — пропустите его, чтобы использовать анонимный бесплатный тариф. Установите WHITEINTEL_API_KEY=wi_…, чтобы авторизоваться по вашему плану и поднять лимиты.
Попробуйте
Вы: «Кто в конечном счёте владеет Revolut? Проверьте санкции по всей цепочке».
Агент: вызывает
search_entities({ query: "Revolut" })→trace_ownership_path({ id })→get_sanctions({ id })для каждого шага → полностью цитируемая цепочка владения с санкционным скринингом на каждом уровне. Готово.
Вы: «Найдите компании, похожие на Wirecard, и проверьте офшорную экспозицию».
Агент: вызывает
find_similar({ entity_id })→check_offshore_exposure({ id })→ выявленные юрисдикции с секретностью и санкционированные посредники по всему набору аналогов.
Агенты могут платить
Агент может купить платную глубину досье от начала до конца, без аккаунта WhiteIntel:
buy_dossier{ tier: "standard" | "premium", entity_id }→ возвращает Stripecheckout_url. Standard (€39) открывает полную многошаговую цепочку UBO + финансовые данные; Premium (€99) добавляет самолёты, санкционированные суда и недвижимость; для субъектов с ВЫСОКИМ риском или санкциями дополнительно запускает живое сканирование негативных публикаций (это сканирование ограничено — оно не выполняется для сущностей с более низким риском).Человек завершает оплату по
checkout_url— Stripe собирает email и перенаправляет обратно.claim_dossier{ session_id }→{ token, entity_id, tier }. Идемпотентно; возвращает402до оплаты.get_dossier{ id, token }→ разблокированное, полностью цитируемое досье в JSON. Токены действительны 90 дней.
Нет человека за клавиатурой прямо сейчас? Шаг 1 — не тот инструмент: checkout_url одноразовый и истекает через 24 часа, так что он уже недействителен к моменту, когда кто-то прочитает ваш отчёт. Вместо этого вызовите get_payment_link — он возвращает постоянные ссылки Stripe, которые можно вставить в документ, тикет или сообщение, и добавьте ?client_reference_id=<entity uuid>, чтобы привязать к конкретной компании. Измерено 2026-08-11: эти ссылки покрывают только тариф Standard (одиночный / 5 / 25); Premium по-прежнему идёт через buy_dossier.
Сначала проверьте get_pricing — он возвращает полный прайс-лист и этот процесс в машиночитаемом виде.
Корпус
~130,7 млн сущностей из 31 объединённого реестра — каждое утверждение цитируется, каждое ребро прослеживается.
Измерено 2026-08-16 по whiteintel.dev/api/public/stats (entities = 130 735 728, само по себе оценка планировщика). Эта конечная точка пересобирает свою карту источников, подсчитывая реестры, поэтому она всегда авторитетна — и новый источник появляется там без правки этого файла.
Источник | Что это | Покрытие |
OpenOwnership | Британские PSC (лица со значительным контролем) | 🇬🇧 Полное |
GLEIF | Глобальный реестр LEI + отношения владения родитель/ребёнок — запланировано ночное обновление | 🌍 Глобальное |
ACRA Singapore | Реестр компаний Сингапура | 🇸🇬 Полное |
ICIJ Offshore Leaks | Panama Papers, Paradise Papers, Pandora Papers | 🌍 Офшорное |
SEC EDGAR | Подачи ценных бумаг США + бенефициарное владение | 🇺🇸 Полное |
UK Companies House | Полный британский реестр — массово + живой поток подач | 🇬🇧 Полное |
FAA | Реестр воздушных судов США (хвостовые номера → владельцы) | 🇺🇸 Полное |
France SIRENE | Французский реестр компаний | 🇫🇷 Полное |
Brazil RFB | Бразильская федеральная налоговая — реестр CNPJ | 🇧🇷 Полное |
Cyprus DRCOR | Кипрский реестр — только должностные лица (см. примечание о границах ниже) | 🇨🇾 Загрузка |
OFAC / EU / UN / UK | Консолидированные санкционные списки | 🌍 Живое |
+ ещё 15 | реестры, санкционные списки и реестры UBO | 🌍 Растущее |
Кипр — что это, а что нет
Кипр вышел в продакшн 2026-08-11 и всё ещё загружается — поэтому мы не приводим здесь замороженное количество строк; спросите /api/public/stats за текущую цифру.
Прочтите это, прежде чем продавать это как покрытие владения Кипром — это не так. Кипрский релиз открытых данных покрывает только номинальный слой: директоров, секретарей и владельцев торговых наименований. Он не содержит ни акционеров, ни бенефициарных владельцев. По выборке загруженных рёбер примерно 93% — это Directorship (директор, секретарь, уполномоченное лицо, генеральный партнёр), а остальные ~7% несут схему Ownership с ролью Owner — это единоличные владельцы торговых наименований, индивидуальный предприниматель, зарегистрированный под коммерческим названием, а не долевое владение компанией. Более ранняя версия этого абзаца утверждала, что в данных Кипра «нет ни одного ребра владения»; это было неверно, и здесь это исправлено, а не тихо удалено, потому что утверждение о том, чего не содержит источник, — это именно та фраза, на которую полагается покупатель.
Практическое следствие не изменилось, и это главное: кипрская компания обычно ответит на trace_ownership_path и check_offshore_exposure значением no_ownership_data. Этот вердикт означает у нас нет рёбер владения для этого субъекта, а не эта компания чисто принадлежит. Не читайте 7% как покрытие акционеров — это не так.
Кипрские записи несут идентификатор cy-reg:. lookup_by_identifier не принимает эту схему — обращайтесь к ним через search_entities с juris: "cy".
Содержит информацию из Департамента регистратора компаний и интеллектуальной собственности Кипра, лицензировано по CC BY 4.0.
Семантический поиск (semantic_search / find_similar) работает по разрешённым карточкам досье с использованием эмбеддингов BGE-M3; покрытие растёт по мере завершения обратного заполнения эмбеддингов. Измерено 2026-08-11 из собственной полезной нагрузки coverage конечной точки: 990 055 из 47 486 969 вселенной встроено (2,1%), и этот срез — ~99,6% в списках риска и ~97% физические лица — так что сегодня эти два инструмента ведут себя гораздо больше как поиск по санкциям/PEP, чем поиск по корпусу, и пустой результат обычно означает «ещё не встроено». Лексический search_entities всегда покрывает весь корпус; используйте его вместе с любым из них, прежде чем делать вывод.
Почему WhiteIntel
Что в имени: White + Intel — white как прозрачный, открытый, цитируемый; intel как разведка, а не данные. Мы не продаём сырые записи — мы продаём разрешение, обход и цитируемую доставку.
Существующие инструменты корпоративного владения были созданы для аналитиков по комплаенсу, кликающих по веб-формам. WhiteIntel — это интеллектуальный слой для эпохи агентов — где исследователем может быть человек, автономный агент или ИИ-рабочий процесс, и всем им нужна одна и та же цитируемая, обойдённая, оценённая по риску разведка.
Цитируется, а не заявляется. Каждое ребро владения, каждый санкционный флаг, каждый сигнал риска прослеживается до записи публичного реестра с реальной датой вступления в силу. Мы не выдумываем и не выводим — если источник этого не говорит, мы этого не говорим.
MCP-нативный, а не ещё один API-обёртка. Семантические интеллектуальные примитивы — а не REST-эндпоинты, втиснутые в определения инструментов. Одна команда, любой MCP-агент.
Freemium по дизайну. Публичный корпус можно свободно исследовать — без входа, без API-ключа, без платного доступа к поиску. Вы платите только за глубину: полные цепочки UBO, слои активов, мониторинг и экспорт.
Агенты могут платить. Единственный MCP-сервер, где агент может исследовать компанию, решить, что ему нужно платное досье, купить его через Stripe Checkout и получить цитируемую разведку — от начала до конца, без человеческого портала.
Честно о пробелах. Отсутствующее ребро означает «ещё не наблюдалось», а не «не существует». Поддержка принятия решений для расследований, а не юридическое определение бенефициарного владения.
Без привязки. Лицензия MIT. Ваш агент, ваши данные, ваше расследование.
Данные и честность
Живой корпус: ~130,7 млн сущностей из 31 объединённого реестра (измерено 2026-08-16). Живые счётчики всегда авторитетнее этого файла: whiteintel.dev/api/public/stats.
Источники не одинаково глубоки. Наличие реестра в списке означает, что мы храним то, что публикует этот реестр — а для некоторых юрисдикций это слой должностных лиц, а не владение. Кипр — самый яркий случай (см. примечание о охвате выше). Никогда не читайте наличие в таблице источников как покрытие владения.
Отсутствующее ребро означает «ещё не наблюдалось», а не «не существует».
Поддержка принятия решений для расследований, а не юридическое определение бенефициарного владения.
Покрытие семантического поиска растёт по мере завершения обратного заполнения эмбеддингов — лексический поиск всегда покрывает весь корпус.
Конфигурация
Переменная окружения | По умолчанию | Назначение |
| (нет) | Необязательный ключ |
|
| Источник API (защищён от SSRF для хостов whiteintel.dev). |
|
| Таймаут на запрос. |
Экосистема
WhiteIntel — часть растущей интеллектуальной платформы:
whiteintel.dev — веб-приложение: поиск, граф, досье, лента Pulse, мониторинг
WhiteIntel API — REST API со спецификацией OpenAPI, те же конечные точки, которые вызывает этот MCP-сервер
WhiteIntel Pulse — живая лента изменений владения/контроля по корпоративному графу
@whiteintel/mcp-server — этот пакет: MCP-слой разведки
Кто за этим стоит
WhiteIntel создан и руководится @Hei33enberg — независимый проект разведки, финансируемый самостоятельно. Никакого венчурного капитала, никаких брокеров данных, никаких компромиссов в целостности цитирования.
Швейцарское управление · Честно по конструкции
Попасть в граф
npx -y @whiteintel/mcp-server # 21 tools, any MCP agentУстановка — поместите сервер в Claude Desktop, Cursor, Cline, Windsurf или в свою собственную среду выполнения (см. Краткое руководство).
Ключ не нужен — работает на анонимном бесплатном уровне из коробки.
Идите глубже — установите
WHITEINTEL_API_KEYдля полной глубины вашего плана.Исследуйте корпус — whiteintel.dev — бесплатный поиск, без входа.
Владейте им — поставьте звезду репозиторию, стройте на API или интегрируйте в свой агентский конвейер. MIT, без привязки.
Вклад
Приветствуются проблемы, PR и идеи инструментов. Начните с CHANGELOG для того, что выпущено и что дальше. Если вы создаёте агента, использующего корпоративную разведку, мы хотим услышать о вас — intel@whiteintel.dev.
Сообщество: GitHub Issues для ошибок и функций, GitHub Discussions для дизайна и помощи.
Веб: whiteintel.dev · npm: @whiteintel/mcp-server · Релизы: GitHub
Лицензия
MIT © whiteintel.dev
Available Tools
21 toolsbuy_dossierAInspect
Start a one-off dossier purchase via guest Stripe Checkout — no WhiteIntel account needed (Stripe collects an email for delivery). Pick a tier ('standard' €39: full UBO chain + financial history · 'premium' €99: additionally itemised assets — vessels, aircraft, securities, real estate) and optionally a bulk pack ('5' or '25' report credits; standard 5×€159 / 25×€599, premium 5×€399 — no premium 25-pack) plus the entity_id (from search_entities) the report is for. Returns checkout_url + next_steps: open the URL so payment can be completed, then feed the session_id from the post-payment redirect to claim_dossier for the access token. See get_pricing for the full price list. WRONG TOOL IF NOBODY IS THERE TO PAY: the session it mints is single-use and expires in 24 hours, so putting this URL in a report or a message read tomorrow hands over a dead link. Use get_payment_link for a permanent, reusable one (standard tier only — Premium is available solely through this tool). And do not fetch checkout_url yourself; it is a card form, so it must be handed to a human.
| Name | Required | Description | Default |
|---|---|---|---|
| pack | No | Optional bulk pack (default single). standard: 5=€159 / 25=€599 · premium: 5=€399 (no 25-pack). | |
| tier | Yes | Dossier tier: standard (€39) or premium (€99, adds itemised assets). | |
| entity_id | No | Optional entity id (from search_entities) the dossier should unlock. | |
| entity_name | No | Optional entity display name, recorded on the Stripe session as an audit trace only — it is NOT displayed anywhere. Since 2026-08-09 the name shown on the invoice and in the delivery email is read from WhiteIntel's own record for entity_id (caller-supplied text is never rendered in mail we send), and the Checkout page shows the Stripe product name. Safe to omit. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses critical behaviors: guest checkout, Stripe collects email, session is single-use and expires in 24 hours, and caller-supplied entity_name is not displayed (safety about audit data). It also warns against fetching the URL. This is comprehensive and goes beyond basics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but efficiently packed with essential information. Each sentence adds value: pricing, required parameters, return value, error conditions, alternatives, and security warnings. Despite length, it is front-loaded with the core purpose and ends with critical usage caveats. No redundant fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though there is no output schema, the description explicitly states the return (checkout_url + next_steps) and the follow-up workflow (feed session_id to claim_dossier). With full parameter documentation and clear behavior, the description is complete for the tool's complexity. It addresses all necessary aspects for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (100%), so baseline is 3, but description adds value: it explains the pricing tiers in prose, clarifies that entity_id comes from search_entities, and details entity_name's audit-only role and its recent behavior change. It supplements the schema with context that helps correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource and clear purpose: 'Start a one-off dossier purchase via guest Stripe доListening surtout mailing类专业 unsaturated бы粱作为一名大笑mesytetiwxx: product name visible anywhere, product name without, "buy_dossier" not mention in "Buy" or , and line: "Start a one-off dossier purchase..." This differentiates from siblings like get_payment_link (permanent link) and claim_dossier (token retrieval). It clearly states the action and the return of checkout_url, and distinguishes from get_pricing and get_payment_link.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool vs alternatives: 'WRONG TOOL IF NOBODY IS THERE TO PAY...' and 'Use get_payment_link for a permanent, reusable one...'. It also warns not to fetch checkout_url programmatically. This provides clear when/when-not guidance and names alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_offshore_exposureAInspect
Walk the ownership chain upward from an entity and flag, hop by hop, whether each node is sanctioned and/or sits in a secrecy jurisdiction (classic tax-haven / offshore-secrecy country). Returns the chain, the flagged hops, and a structured 4-state verdict — BRANCH ON verdict, NOT on exposed. States: no_ownership_data (we hold zero ownership edges from this entity — NOT a clean verdict, exposure cannot be evaluated), flagged (a sanctioned or secrecy-jurisdiction hit sits on the walked chain), checked_to_max_depth_truncated (walk reached the depth cap with more chain above — a flagged owner may still sit higher, NOT clean), checked_full_clean (the walk ran out of chain before the cap, no flag). Also returns depth_walked (how deep the walk actually reached) and depth_capped. READ depth_capped EVEN WHEN THE VERDICT IS checked_full_clean, because the two co-occur. Measured 2026-08-11 anonymously with max_depth=6: verdict: 'checked_full_clean', depth_walked: 1, depth_capped: true, plan: 'free'. depth_capped: true means A CAP WAS IN FORCE, not that the cap necessarily bit — here the chain genuinely ended after one hop, below the free plan's 2-hop ceiling. The honest report of that response is 'clean over the one hop of ownership we hold, on a walk a free key limits to two', which is what the payload's own note says in prose. Never promote checked_full_clean to 'no offshore exposure' without quoting depth_walked. Anonymous callers walk at most 2 hops however high you set max_depth. Legacy exposed boolean is retained but is only meaningful when verdict='flagged'. Get the id from search_entities or lookup_by_identifier.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id to assess. | |
| max_depth | No | Max ownership hops to request (default 6). Anonymous callers are capped at 2 — read `depth_walked` in the response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It details depth caps, the meaning of depth_capped, the legacy exposed field's limited usefulness, and a concrete anonymous-caller example, giving the agent excellent expectations for edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but the complexity of the verdict semantics justifies much of the length. It is front-loaded with the core purpose and then walks through states and caveats; however, the measured example and repeated warnings could be tightened without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description thoroughly explains return values: the chain, flagged hops, verdict states, depth_walked, depth_capped, and legacy exposed. It also covers cap behavior and id sourcing, making the tool usable without additional external documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema covers both parameters, the description adds meaningful semantic value: it tells the caller to obtain id via search_entities or lookup_by_identifier, and explains that anonymous callers are capped at 2 hops regardless of max_depth. This goes beyond the schema's basic field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Walk the ownership chain upward from an entity and flag...', which clearly identifies what the tool does. It also distinguishes this tool from siblings like trace_ownership_path or get_sanctions by emphasizing the offshore/sanctions-secrecy verdict semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong interpretive guidance, such as 'BRANCH ON verdict, NOT on exposed', and warns against promoting checked_full_clean to 'no offshore exposure' without quoting depth_walked. It does not explicitly name alternatives or exclusions, but the context for appropriate use is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claim_dossierAInspect
Redeem a paid Stripe Checkout session for a dossier access token. Pass the session_id (cs_…) from the post-payment redirect after buy_dossier. Returns { token, entity_id, tier } — pass the token to get_dossier as its token input for the unlocked report (standard: full UBO chain + financial history · premium: additionally itemised assets). Idempotent: claiming the same session again returns the same grant, so it is safe to retry. Fails with 402 not_paid until the payment has actually completed — wait for the human to finish Checkout, then call again.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Stripe Checkout session id (cs_…) from the success redirect. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description fully carries the burden of behavioral disclosure. It reveals idempotency (same session returns same grant, safe to retry), handles failure (402 not_paid until payment completed), and explains the return structure and tier variations. This is comprehensive and directly addresses operational behaviors an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured paragraph. It front-loads the core purpose, then provides usage, return details, and error handling without unnecessary verbosity. Every sentence adds value, no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description thoroughly explains return values ({ token, entity_id, tier }) and how to use them. It also covers idempotency, retry safety, and error conditions. With no annotations, this is as complete as one could expect for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage for session_id with a clear description. The description adds contextual meaning by explaining where to get the session_id (from the post-payment redirect after buy_dossier) and reiterating the format (cs_...). This goes slightly beyond the schema, so it earns a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: redeem a paid Stripe Checkout session for a dossier access token. It uses a specific verb+resource (claim dossier) and distinguishes itself from sibling tools by referencing buy_dossier and get_dossier, which is exactly the kind of differentiation needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: after the post-payment redirect following buy_dossier. It also tells the agent to pass the returned token to get_dossier, and explains when not to call (before payment completion), including the 402 not_paid failure and the retry guidance. This is explicit usage context with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_similarAInspect
Entities most similar to a given one — the nearest corpus dossier cards ('more like this'), for peer discovery and clustering around a known entity. Pass an entity_id from search_entities. Returns { id, count, hits }, each hit with entity_id, caption, kind, jurisdiction, risk and a similarity score. COVERAGE IS PARTIAL AND SKEWED — it draws on the same embedded slice as semantic_search: 990,055 of a 47,486,969 universe (2.1%), ~99.6% risk-listed and ~97% natural persons, measured 2026-08-11 from the sibling endpoint's own coverage payload. An entity outside that slice returns count: 0 with an empty hits array and HTTP 200 — that is 'not embedded', NOT 'no peers exist', and it is the common case for ordinary companies (verified: BARCLAYS BANK PLC returns zero). Never report an empty result as a finding about the entity. Fall back to semantic_search or search_entities.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Max hits (default 10). | |
| entity_id | Yes | Entity uuid from search_entities. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and excels: discloses partial coverage (990,055/47,486,969 = 2.1%), skew (99.6% risk-listed, 97% natural persons), behavioral nuance (HTTP 200 with count:0 for non-embedded entities), and a verified example (BARCLAYS BANK PLC). Even warns against misinterpreting empty results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place. The description front-loads the core purpose, then returns format, then critical limitation warnings. It is longer than average but all content is necessary behavioral caveats, and it remains well-structured and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates by detailing the return shape ({ id, count, hits } with per-hit fields). It covers the tool's purpose, limitations, fallback alternatives, and edge-case behavior, making it fully self-contained for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. Description reinforces that entity_id must come from search_entities, but adds little beyond the schema's existing descriptions for entity_id and k. It does not explain k's effect beyond schema, but no compensation needed due to high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: 'Entities most similar to a given one — the nearest corpus dossier cards ('more like this')'. It explicitly names the intended use case (peer discovery, clustering) and distinguishes itself from semantic_search and search_entities via fallback guidance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use: 'for peer discovery and clustering around a known entity' and clear exclusion: coverage is partial/skewed, empty results mean 'not embedded' not 'no peers exist'. Directly names alternatives: 'Fall back to semantic_search or search_entities.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_company_detailsAInspect
Companies House register detail for a UK company by entity id: registered address, status, company type, incorporation date, SIC industry codes, and the filing/compliance layer — accounts type, last-filed and next-due dates (flagged when OVERDUE), confirmation-statement status, outstanding mortgage charges, and former ('also known as') names. Use this for 'where is X registered / what does it file / is it overdue / what was it called before'. Returns { entity, company_details, provenance, note, source } — this is the best-populated of the UK detail tools, measured 2026-08-11 at 45 of 48 sampled UK company entities carrying a non-empty company_details (contrast get_financials at 11 of the same 48). Get the id from search_entities or lookup_by_identifier.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id (a UK company). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return structure ('Returns `{ entity, company_details, provenance, note, source }`'), notes data completeness with a specific measurement date and sample size, and flags overdue statuses. It does not mention potential errors, rate limits, or auth requirements, but for a read-only lookup tool, the provided behavioral context is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense but well-structured, front-loading the core purpose and data fields, then usage guidance, return format, and data-quality metric. It is slightly long but every sentence adds value, including the comparative metric against get_financials. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (single parameter, no output schema), the description is quite complete: it lists the data fields, return envelope, usage examples, and data-quality context. It could mention pagination or error behavior, but for a single-id lookup with no output schema, the description covers the essential context well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage for the single parameter 'id' with a description ('Entity id (a UK company).'). The description adds that the id is for a UK company and how to obtain it, but does not add significant new semantics beyond the schema. Baseline 3 is appropriate given full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves Companies House register details for a UK company by entity id, listing specific data fields (registered address, status, company type, incorporation date, SIC codes, filing/compliance layer). It distinguishes itself from siblings by noting it is the 'best-populated of the UK detail tools' and contrasts with get_financials, making its purpose and differentiation explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: 'Use this for 'where is X registered / what does it file / is it overdue / what was it called before''. It also tells the agent how to obtain the required id ('Get the id from search_entities or lookup_by_identifier'), and implicitly contrasts with get_financials for financial data, giving clear when-to-use context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dossierAInspect
Build a structured, fully-cited intelligence dossier for one entity by id: identity with cross-source linked records (the same real-world entity resolved across ICIJ leaks, GLEIF, registries), ownership/control (direct owners, holdings, and the UBO chain), risk signals, and provenance on every layer. Every claim traces to a source URL. Use this for 'tell me everything about X'. Get the id from search_entities. Free tier shows the first ownership hop + latest financials; pass a one-off purchase token (from claim_dossier, see get_pricing / buy_dossier) or set WHITEINTEL_API_KEY to unlock the full depth.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id (from search_entities). | |
| token | No | Optional one-off dossier access token (from claim_dossier or the delivery email). A standard token unlocks the full UBO chain + financial history for this entity; a premium token additionally unlocks itemised assets. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It does so by disclosing the free tier limitation ('first ownership hop + latest financials'), the token/API key unlock mechanism, and provenance guarantees ('Every claim traces to a source URL'). It does not cover rate limits or error handling, but provides strong contextual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a solid paragraph without fluff. Each sentence contributes value: purpose, provenance, use case, id source, and tier/token details. It could be slightly shortened, but it is appropriately sized for a complex paid tool with no annotations or output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex dossier tool with no annotations and no output schema, the description is remarkably complete. It explains what the dossier includes (cross-source records, ownership/control, UBO chain, risk signals), what the user gets on free tier, and how to unlock full depth. It doesn't describe the exact return format, but 'structured dossier' implies sufficient structure for the agent to proceed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds context about where the id comes from (search_entities) and where the token comes from (claim_dossier), plus the environment variable alternative. However, the schema already fully documents both parameters, and the description adds little beyond acquisition channels.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Build a structured, fully-cited intelligence dossier for one entity by id'. It clearly distinguishes this from siblings like search_entities ('Get the id from search_entities') and get_entity by emphasizing the comprehensive dossier nature and the 'tell me everything about X' use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit use case ('Use this for "tell me everything about X"') and a prerequisite ('Get the id from search_entities'). It also explains free vs paid tiers and token acquisition, but does not explicitly name alternative tools or say when NOT to use it, which is a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_entityAInspect
Full record for one entity by id: type (company/person), identifiers, jurisdiction, risk level, summary and its direct relationships with provenance. Get the id from search_entities or lookup_company.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It discloses the main output shape, listing supported fields, and clarifies the entity types. However, it does not mention caveats, limits on relation depth, formatting, authentication needs, or what happens if the id does not exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and pack is highly relevant. Every phrase adds something meaningful: the action, the input, the output fields, and the id-source relationship.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description is reasonably complete: it states input, output fields, entity types, provenance, and how to get the id. It could be improved by briefly mentioning what is not included or why this is distinct from get_dossier/get_company_details, but the essential usage is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only documents 'Entity id' but the description adds useful semantic guidance: the id refers to a specific entity full record, and the id is typically derived from search_entities or lookup_company. This helps agents understand what value to pass and how to obtain it without duplicating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description explicitly names the tool's action ('Full record for one entity by id') and enumerates the returned fields (type, identifiers, jurisdiction, risk level, summary, direct relationships with provenance). It distinguishes itself from search tools by referring to them as id-sources, establishing a clear get-by-id purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides direct workflow guidance: 'Get the id from search_entities or lookup_company.' This clearly implies the primary use case (fetch a full entity record once id is known), but it does not list explicit alternatives or when-not-to-use scenarios relative to other sibling tools like get_dossier or get_company_details.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_financialsAInspect
Filed financial figures for a UK company by entity id, year-over-year, from Companies House iXBRL accounts: turnover, profit/(loss), net assets, cash, shareholder funds, fixed/current assets, and employee count per reporting period. Use this for 'what are X's revenue / profit / net assets / how many employees'. Returns { entity, financials, note, source }. MOST ENTITIES HAVE NOTHING HERE, AND THAT IS THE NORMAL ANSWER, NOT AN ERROR. Measured 2026-08-11 on a sample of 48 UK company entities drawn from search_entities: only 11 returned any filed period — the other 37 came back HTTP 200 with an empty financials and a note saying so (even BARCLAYS BANK PLC, CH 01026167, has none loaded). Earlier versions of this description called balance-sheet coverage 'broad'; it is not. Within the accounts we DO hold, the per-field skew is real: balance-sheet items and employee counts are the well-populated ones, while turnover and profit are sparse because micro-entities file no profit-and-loss account. Read note before writing 'no revenue' — absent filings and a filed zero are different claims. Get the id from search_entities.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id (a UK company). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden and exceeds it: it discloses the source, the exact return shape, that most entities return empty financials with a note (with a concrete measured sample and even a Barclays example), that balance-sheet fields are well-populated while turnover/profit are sparse, and that absent filings must not be conflated with filed zeros. This is unusually rich operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence earns its place: purpose, use-case, return shape, normal-empty caveat, measured evidence, field-skew warning, and note-reading instruction. It is front-loaded with the core function and flows logically from purpose to usage to interpretation, with no filler or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by stating the return structure, enumerating the reported financial fields, explaining the empty-financials behavior, and clarifying the meaning of the note field. For a tool whose main risk is misinterpreting missing data, this is a complete and robust description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes id as 'Entity id (a UK company)' at 100% coverage, so this dimension starts at baseline 3. The description adds meaningful provenance ('Get the id from search_entities') and reinforces that the id must be a UK company entity id, going slightly beyond the schema without needing more detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Filed financial figures for a UK company by entity id' using Companies House iXBRL accounts, and enumerates the exact metrics (turnover, profit/loss, net assets, cash, shareholder funds, fixed/current assets, employee count). This clearly distinguishes get_financials from sibling tools such as get_dossier or get_company_details, which serve broader company-profile purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit 'use this for' clause ('what are X's revenue / profit / net assets / how many employees') and points to search_entities as the source for the id. It does not explicitly name alternative tools for other data types or state a 'when not to use' condition, but the coverage warning strongly informs expectation-setting around use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_payment_linkAInspect
PERMANENT, shareable Stripe payment links for the one-off dossiers — use this INSTEAD of buy_dossier whenever you need something you can HAND TO A HUMAN. buy_dossier mints a Checkout Session that is single-use and expires in 24 hours, so it is useless in a report, a ticket or a message the human reads tomorrow; these links never expire and can be reused. Append ?client_reference_id= to bind the purchase to one company — without it the buyer gets a dossier credit, spendable on any entity later. No API key and no WhiteIntel account needed. MEASURED 2026-08-11: the response carries STANDARD-tier links only — single (€39), 5-pack (€159) and 25-pack (€599). There is no Premium payment link, so for Premium (€99) you must still use buy_dossier and have someone finish Checkout inside 24h. You cannot complete any of these yourself: the page is a card form.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description takes on a substantial information burden. It explains link permanence, reusability, the lack of need for an API key or WhiteIntel account, and the ability to bind to a company via client_reference_id. It also discloses the exact response tier structure and its limitation (no Premium link).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, using bolding and uppercase for emphasis. While it's long, every sentence provides necessary context for a financial transaction tool. The most critical distinguishing points (permanent, shareable) appear first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema, this description is complete. It covers purpose, usage restrictions, behavioral traits, pricing structure, and limitations. It fully compensates for the absence of structured annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is minimal and 100% covered. The description even adds the optional client_reference_id query parameter context, though that isn't an explicit input. Given 0 params, baseline is 4, but the detailed context around query string binding elevates this to 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies this as the tool for permanent, shareable Stripe payment links for one-off dossiers. It emphasizes the distinction from buy_dossier (single-use, 24h expiry), which is the key to its purpose, explicitly naming the sibling tool to prevent confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly contrasts with buy_dossier, stating 'use this INSTEAD of buy_dossier whenever you need something you can HAND TO A HUMAN.' It also explains when buy_dossier is required (for Premium), thereby offering clear alternatives and exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pricingAInspect
WhiteIntel's price list plus the exact machine flow for buying access. One-off cited dossiers (Standard €39: full UBO chain + financial history · Premium €99: additionally itemised assets), bulk packs (5× / 25× at a discount), subscriptions (Investigator €149/seat·mo, Business €1,900/mo) and the metered API. Returns how_an_agent_buys — buy_dossier opens a Stripe Checkout, a human (or payment-capable agent) pays, claim_dossier mints the access token, and get_dossier with that token returns the unlocked report. Step 0 of that list covers the case with no human present: get_payment_link returns permanent Stripe links you can hand over instead. Static data, no network call — check it before recommending a purchase.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and delivers: 'Static data, no network call' explicitly declares the behavioral contract, alerting the agent that no side effects or costs are involved. It also discloses return semantics ('Returns how_an_agent_buys') and reveals the purchase-side behavior requiring external payment ('buy_dossier opens a Stripe Checkout, a human... pays, claim_dossier mints the access token').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though long, the description front-loads the key purpose first, then organizes information in digestible parentheticals and enumerations without redundancy. Every element — tiers, discounts, flow, edge case, and behavioral flag — earns its place; the final 'Static data, no network call — check it before recommending a purchase' is a dense, purposeful closer with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 params, no annotations, no output schema), the description is remarkably complete: it covers all pricing products, describes the full order-of-operations among siblings, handles the human-less edge case, and flags its read-only network-free nature. There are no gaps an agent would need clarified to use this appropriately when recommending a purchase.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters with 100% trivial schema coverage, so per the rubric baseline is 4. The description enhances this by specifying what the return payload contains (pricing tiers per product line, discounts, and the purchase flow walkthrough), going beyond the bare schema by describing the data an agent would consume. No parameter documentation burden exists here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause, 'WhiteIntel's price list plus the exact machine flow for buying access,' uses a specific verb-plus-resource phrasing that unmistakably identifies what the tool returns. It also distinguishes itself from siblings by clarifying this is static reference data explaining the purchase flow rather than the buying action itself — the description explicitly differentiates from buy_dossier, claim_dossier, get_dossier, and get_payment_link.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance with 'check it before recommending a purchase' and walks through the full purchase pipeline, including an alternative with a conditional discriminator: 'Step 0 of that list covers the case with no human present: get_payment_link returns permanent Stripe links you can hand over instead.' It even notes the metered API tier and the no-human-present edge case, giving the agent a decision map of when each sibling applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pulseAInspect
The WhiteIntel Pulse activity feed: recent corpus events — ownership/control changes, filed accounts, watchlist and sanctions designations — newest first, each with a source registry. Use this to answer 'what changed recently' or to monitor the corpus. All four kinds are live and all four are cited: measured 2026-08-11, 100 of 100 rows carried a source_url for every kind. THE UNFILTERED FEED IS NOT A BALANCED MIX. It applies no kind filter, but it is ordered by ingest recency, so whichever loader ran last fills the head of it. Measured 2026-08-11: the newest 100 rows of the default feed were 100% kind: 'ownership', all from one registry. If you want a particular stream — or any sense of proportion between them — pass kind and do not infer 'nothing else happened' from the default page.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Optional: filter by event kind. Registries measured on the newest 100 rows of each stream, 2026-08-11 — they move as loaders run, so read each row's `source_registry` rather than trusting this note. `ownership` = registry-recorded control changes; today 100/100 came from `borme` (Spain's Boletín Oficial del Registro Mercantil), NOT GLEIF as earlier versions of this description claimed. `filing` = UK Companies House accounts (100/100 `companies_house`). `watchlist` = OpenSanctions non-sanctions listings — politically-exposed persons, criminal/wanted entries and procurement debarments, not PEPs alone (100/100 `opensanctions`). `sanction` = a designation landing on a sanctions list, e.g. OFAC SDN (100/100 `opensanctions`). | |
| limit | No | Max events (default 40). | |
| since | No | Optional sync cursor (ISO-8601): pass the next_since from your last response to get only events ingested after it — poll this to monitor what's new. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to lean on, the description carries the full burden and delivers extraordinary disclosure: the feed is 'NOT A BALANCED MIX', it is 'ordered by ingest recency', and it reveals a measurement that the newest 100 rows were '100% kind: ownership, all from one registry' (2026-08-11). It even corrects prior misinformation ('NOT GLEIF as earlier versions of this description claimed') — the kind of bias disclosure that would be invisible to an agent without this text.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core definition, then usage, then warnings — good structure with strong leading content. However, the measurement statistics (100/100, dates, percentages) appear twice: once in the main prose and again inside the `kind` enum descriptions, creating slight redundancy. Every sentence otherwise earns its place; the uppercase emphasis is effective but the near-duplicate measurement reports could be consolidated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema and no annotations, the description fully compensates: it covers what is returned (events with a source registry), ordering semantics, parameter behavior, per-kind meaning, and dangerous default-bias caveats. For a monitoring/feed tool of moderate complexity, there is nothing essential an agent needs to know that the description omits — it even notes when defaults 'move as loaders run,' setting correct expectations about non-determinism.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description goes far beyond, defining each enum value with measured composition — e.g., 'watchlist = OpenSanctions non-sanctions listings — politically-exposed persons, criminal/wanted entries and procurement debarments, not PEPs alone' — and clarifies the registry source (borme/companies_house/opensanctions). `since` is meaningfully framed as a sync cursor for polling rather than a plain date filter. This is exactly the value the schema's one-line hints forfeit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a crisp noun phrase identifying the exact resource and behavior: 'The WhiteIntel Pulse activity feed: recent corpus events... newest first, each with a source registry.' It enumerates the four event kinds, states the ordering, and notes provenance — a specific verb+resource that is impossible to confuse with the sibling search/lookup/graph tools despite no sibling being named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-to-use: 'Use this to answer "what changed recently" or to monitor the corpus.' It also supplies when-not-to-infer guidance ('do not infer "nothing else happened" from the default page') and instructs the agent to pass `kind` when it wants a particular stream. This directly helps the agent choose between this and the search/entity siblings, even without naming them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sanctionsAInspect
Return an entity's screening exposure for the entity AND its resolved cluster siblings, each with a source URL. IT IS NOT SANCTIONS-ONLY, DESPITE THE NAME — read each row's signal_type. Measured 2026-08-11: BARCLAYS BANK PLC came back sanctioned: false with one signal of signal_type: 'crime' (severity HIGH, source_list opensanctions_crime, a criminal/wanted listing reaching it via its cluster). Only signal_type: 'sanctioned' rows are sanctions designations, and only those reliably carry list and regime — on the crime row both were null, so do not read a null list as missing data. Two consequences: a sanctioned: false response can still contain a HIGH-severity adverse finding you must report, and 'no sanctions signal' (what the top-level flag and note describe) is not 'nothing found'. Response splits the top-level flag: sanctioned_self = a direct listing ON this entity; sanctioned_via_cluster = the flag reaches it ONLY via a cross-source cluster sibling (~2.3% false-positive tail on UK OpenOwnership resolution — treat cluster-only hits as a lead until you verify the sibling really is the same real-world party). The aggregate sanctioned (self OR cluster) is preserved for back-compat. Get the id from search_entities or lookup_by_identifier.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Entity id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals cluster expansion, signal_type semantics, null list/regime behavior, the distinction between sanctioned_self and sanctioned_via_cluster, the false-positive tail, and back-compat behavior. This is far beyond what annotations would have provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: it front-loads the core function, then delivers critical caveats, a concrete measured example, flag semantics, and id-source guidance. The structure is logical and dense without redundancy, which is appropriate given the tool's misleading name and nuanced behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the absence of an output schema, and the absence of annotations, the description is remarkably complete. It covers return contents, signal types, null handling, flag meanings, false-positive risk, and how to obtain the required id. An agent has enough information to invoke the tool and interpret its results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the single `id` parameter with 100% coverage. The description adds useful semantic value by specifying that the id should come from search_entities or lookup_by_identifier, and by framing the id as an entity id. This goes beyond the schema's bare 'Entity id' description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Return an entity's screening exposure for the entity AND its resolved cluster siblings, each with a source URL.' It also explicitly distinguishes the tool from its misleading name by clarifying it is not sanctions-only and by directing attention to `signal_type`. This clearly separates it from sibling tools like get_entity or lookup_by_identifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong contextual guidance: it explains what the tool returns, warns that a `sanctioned: false` response can still contain adverse findings, and tells the agent to get the id from search_entities or lookup_by_identifier. It does not explicitly enumerate when not to use the tool versus alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_neighbourhoodAInspect
Return every ownership/control edge within a bounded number of hops of one entity, in BOTH directions: who it controls, who controls it, and their neighbours. Use it to answer 'what sits around this company?' — the wider view that trace_ownership_path (upward only) does not give. Hard-capped in the database: depth 3, 300 edges, and at most 25 edges followed per entity per direction per hop. READ THE DEPTH FIELDS IN THE RESPONSE — DO NOT ASSUME YOU GOT THE DEPTH YOU ASKED FOR. There is no field called depth any more, and that rename is deliberate: the old depth was the CLAMPED REQUEST, never the depth walked, and it was being read as a promise. The response now carries depth_requested (what your plan allowed), depth_walked (measured off the returned edges' own hop numbers — the only depth that is actually proven), depth_capped, and completeness. Measured 2026-08-11 on an anonymous caller: depth=3 requested returned depth=2 with depth_capped=true, because the free plan caps every walk at 2 hops. Any sentence you write about what is or is not around this entity must be scoped to the RETURNED depth. EDGE COUNTS FELL BY UP TO 2.7x ON 2026-08-11 AND NOTHING WAS LOST — read this before you treat it as the corpus shrinking. Until that date the walk emitted the same edge two and three times at depth 2 or more, edge_count counted the duplicated list, and the duplicates were charged against your edges budget. Measured on identical requests before and after the fix: 72 -> 27, 29 -> 13, and at the maximum budget 300 rows holding 285 real edges -> 300 rows holding 300. So a call you made yesterday and repeat today can return far fewer edges for the same subject: the smaller number is the true one, and your budget now buys real edges. One consequence worth knowing: at depth 1 a root can drop from 4 edges to 2, because the registry genuinely holds rows that are identical in every field this endpoint returns and the response has no way to represent the difference. That is also a correction, not a loss. truncated: true plus a plain-language truncation_note does work and does mean the edge budget ran out (verified with edges=10); that is NORMAL for hub entities (the corpus holds single nodes with more than 22,000 edges) and means the picture is partial, not wrong. Each edge carries origin: 'registry' (observed in a source registry) or 'derived'/'curated'/'asserted' (inferred by WhiteIntel). Get the root id from search_entities or resolve.
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | Root entity uuid. | |
| depth | No | Hops to walk (default 2). A REQUEST, not a guarantee — the plan caps it (anonymous callers measured at 2 hops) and the response's `depth_walked` is the authority — it is measured from the edges that came back, not echoed from your request. | |
| edges | No | Edge budget (default 120). Lower it for a legible picture, raise it for completeness. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full weight. It discloses hard caps, depth clamping, deduplication changes, truncation behavior, edge-count expectations, and a warning not to trust requested depth. This is far more transparent than typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and sibling distinction, which is helpful. However, it becomes quite long and mixes historical measurements, explicit troubleshooting notes, and warnings in a way that requires careful parsing; a tighter summary would improve scanability without losing key caveats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description provides important response semantics: depth_requested, depth_walked, depth_capped, completeness, truncation, edge_count, and edge origin. It also covers edge cases around hubs and duplicate records, making the tool safe to invoke even under complex conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers 100% of parameters, so the baseline is 3. The description adds meaningful semantic context for the `depth` and `edges` parameters (request vs actual depth, budget and deduplication) and points to source endpoints for `root`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states explicitly that it returns ownership/control edges within a bounded number of hops in both directions, with entity neighbours. It also distinguishes itself from the sibling `trace_ownership_path` by noting that tool is upward-only, making the tool's purpose and scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use case: 'what sits around this company?' and explicitly contrasts the wider neighbourhood view with `trace_ownership_path` (upward only). It also tells the agent how to obtain the root id via `search_entities` or `resolve`, so invocation prerequisites are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_pathAInspect
Find how two entities are connected: a bounded breadth-first search over ownership and control edges in both directions, returning the ordered hops from one to the other. WARNING, AND IT CHANGES HOW YOU MUST REPORT THE RESULT: this search is BOUNDED, NOT EXHAUSTIVE. At most 15 edges are followed per entity, per direction, per hop, so a genuine connection running through a heavily-connected intermediary can be missed. found: false means NO PATH WAS FOUND WITHIN THOSE BOUNDS — it is NOT evidence that the two entities are unconnected, and must never be reported as a clean result. The response always carries exhaustive: false, a structured verdict (e.g. 'connected_within_bounds') and a bounds_note restating this. AND THE DEPTH YOU GET IS NOT THE DEPTH YOU ASK FOR: the response echoes its own max_depth plus depth_capped, and those are the authority. Measured 2026-08-11 anonymously — max_depth=3 and max_depth=4 both came back as max_depth: 2, depth_capped: true, plan: 'free'. So a free-tier found: false is a two-hop negative however many hops you requested; say two hops, not four.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | End entity uuid. | |
| from | Yes | Start entity uuid. | |
| max_depth | No | Max hops to request (default 3). Reduced by the plan — anonymous callers measured at 2 — so read the response's `max_depth` and `depth_capped`. Depth 4 is measurably slower on densely connected entities; request it deliberately. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It excels here, detailing the bounded BFS, non-exhaustive nature, depth-capping behavior, free-tier plan limitation, and the meaning of found:false. It also discusses response fields like exhaustive, verdict, bounds_note, and depth_capped, giving the agent comprehensive insight into runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose and heavily repetitive, with multiple ALL-CAPS warnings and run-on sentences. While every sentence adds some information, the structure is not concise and would benefit from tighter organization. The core message could be delivered in half the length without losing impact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool and the lack of annotations or output schema, the description is highly complete. It covers edge cases (bounded search), failure semantics (found:false), response structure (verdict, bounds_note, depth_capped), plan limitations, and performance characteristics, ensuring the agent is well-informed about all important behaviors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already documents all parameters, so the baseline is 3. The description notably enhances understanding of max_depth by warning that the response's max_depth may be reduced and advising deliberate use of depth 4. However, it adds no new meaning for the from/to parameters beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Find how two entities are connected' via a bounded breadth-first search over ownership and control edges, returning ordered hops. This is a specific verb+resource description, but it does not explicitly distinguish itself from sibling tools like trace_ownership_path or graph_neighbourhood.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While the description provides extensive operational warnings (e.g., bounded search, not to report found:false as clean), it offers no explicit guidance on when to choose this tool over siblings or what alternatives exist. The usage context is implied but no exclusions or comparisons are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_by_identifierAInspect
Resolve an entity by a strong external identifier instead of a name — a LEI, OFAC SDN uid, EU/UN/UK sanctions id, Singapore UEN, SEC CIK, Polish KRS, UK Companies House number, French SIREN, or Brazil RFB CNPJ. Returns the single resolved entity (id, type, jurisdiction, identifier, risk) so you can pivot into get_entity / get_dossier / get_sanctions. Use this when you already hold a registry id and want the corpus node behind it. All eleven schemes were exercised against production on 2026-08-11 and every one resolved a real entity — no scheme in this enum is decorative. DISTINGUISH THE TWO FAILURE MODES: an unsupported scheme returns HTTP 400 with error: 'bad_request' and the accepted set spelled out in detail, whereas a supported scheme whose value we simply do not hold returns HTTP 404 error: 'not_found'. A 404 is a statement about the corpus, not about the tool — fall back to search_entities. NOT every identifier you may see in a response is resolvable here — the enum below is the complete accepted set and the route hard-rejects anything else with a 400. In particular Cyprus records carry a cy-reg: identifier that this tool does NOT accept (verified: cy-reg → 400), and neither is the cusip: seen on US securities rows: reach Cypriot companies with search_entities using juris='cy'.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | The identifier value (e.g. an LEI, an OFAC SDN uid, a Companies House number). | |
| scheme | Yes | Identifier scheme: lei | ofac | eu | un | uk | uen | sec | krs | gb-coh | siren (French SIREN, 9 digits) | br-cnpj (Brazil RFB CNPJ; accepts 8-digit root or full 14-digit form 12.345.678/0001-95). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses return shape (single entity with id, type, jurisdiction, identifier, risk), distinguishes two failure modes (400 for unsupported scheme, 404 for not found), and states that all schemes were production-tested. It also explicitly documents validation behavior (complete enum, hard-reject with 400).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, usage, failure modes, and exclusions are each addressed. It is front-loaded with the core purpose and structured logically, avoiding redundancy despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 schemes, two error modes, non-accepted identifiers) and no output schema, the description is remarkably complete. It covers return value shape, error handling, fallback guidance, and test validation date, leaving no major gap for an agent to misinterpret.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by explaining the complete accepted set, giving examples of non-accepted identifiers, and detailing error responses tied to parameters. However, much of the parameter detail (e.g., br-cnpj formats) is already in the schema, so the incremental addition is moderate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it resolves an entity by strong external identifier, lists specific identifier types, and explicitly distinguishes from search_entities as fallback. The verb 'resolve' and resource 'entity' are specific, and it scopes to single entity results, differentiating it from sibling tools like lookup_company which uses name-based lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('when you already hold a registry id'), when-not-to-use with fallback ('A 404 is a statement about the corpus, not about the tool — fall back to search_entities'), and names alternatives (search_entities for Cyprus ids). Also excludes unsupported identifiers (cy-reg, cusip) with concrete advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_companyAInspect
Look up a UK company by its Companies House registration number and return the company record plus a ready-built ownership graph (officers, persons of significant control, parent/subsidiary edges). Pass the number verbatim — do not strip leading zeros (e.g. 09446231, SC123456).
| Name | Required | Description | Default |
|---|---|---|---|
| number | Yes | UK Companies House registration number. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the disclosure burden. It clearly states the input constraint behavior (verbatim number, preserve leading zeros) and the return behavior (company record plus ownership graph with specified edge types). It does not discuss missing-company behavior, but for a lookup tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the main action, then provides the key nuance (leading zeros) and representative examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one required parameter, full schema coverage, and no output schema, the description provides a clear picture of both input and output expectations. It names the graph components, making the tool sufficiently complete for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers the single parameter, so the baseline is 3. The description adds meaningful semantics beyond the schema: the number must be passed verbatim, leading zeros must not be stripped, and concrete examples are given. This helps the agent handle real-world Companies House identifiers correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Look up a UK company by its Companies House registration number.' It also distinguishes itself from siblings by explicitly promising a 'ready-built ownership graph' (officers, PSCs, parent/subsidiary edges), which differentiates it from search_companies or get_company_details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the primary use case obvious: pass an exact Companies House registration number. It gives clear input guidance ('Pass the number verbatim — do not strip leading zeros') with examples. It does not explicitly mention when a sibling like search_companies should be used instead, but the context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolveAInspect
Batch-resolve a list of company names or strong identifiers (scheme:value — lei, siren, gb-coh, uen, br-cnpj, sec, ofac, eu, un, uk, krs) to canonical WhiteIntel entity ids in ONE call. Each result carries a confidence: 'exact' (identifier match) or 'name' (top name hit); an unmatched row comes back as { match: null, confidence: null }, so check for it rather than assuming positional success. Use this to enrich a whole list — suppliers, counterparties, a portfolio — without one lookup per row. Then feed the ids into get_dossier / trace_ownership_path / get_sanctions. Up to 25 items anonymously (a 26th returns HTTP 400 with the limit spelled out), 100 with WHITEINTEL_API_KEY. TREAT confidence: 'name' AS A CANDIDATE, NOT A RESOLUTION. It is the top lexical hit and nothing more — measured 2026-08-11, the query 'Tesco' resolved to a FRENCH company literally named TESCO (fr-siren:454067281), not Tesco PLC, while 'gb-coh:00445790' resolved 'exact' to TESCO PLC. Confirm a 'name' match's jurisdiction and identifier before you attach it to a real counterparty; pass an identifier whenever you hold one.
| Name | Required | Description | Default |
|---|---|---|---|
| queries | Yes | Names or scheme:value identifiers, e.g. ["Tesco", "siren:552081317", "lei:213800...", "gb-coh:00445790"]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses behavior thoroughly: it explains confidence levels ('exact' vs 'name'), handling of unmatched rows (returns `{ match: null, confidence: null }`), and warns that 'name' matches are candidate hits, illustrated with a concrete Tesco example. This transparency is essential since no annotations are provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is excessively verbose and repetitive. Key information (e.g., the Tesco example, the confidence semantics, and the limits) is stated multiple times, significantly bloating the text. It could be condensed to a few sentences without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides comprehensive context, including usage scenarios, limits, and edge cases. However, the completeness is marred by redundancy; while all necessary details are present, the over-explanation detracts from efficiency. A more concise version would achieve the same completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the 'queries' parameter in detail: it accepts a list of company names or strong identifiers in 'scheme:value' format, with examples like 'siren:552081317' and 'gb-coh:00445790'. This fully clarifies the expected input structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: batch-resolving a list of company names or strong identifiers to canonical WhiteIntel entity IDs. It also mentions the output format and confidence levels, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises when to use the tool: 'Use this to enrich a whole list — suppliers, counterparties, a portfolio — without one lookup per row.' It also notes the batch size limits (25 anonymously, 100 with API key), providing clear operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_companiesAInspect
Free-text company-name search against UK Companies House. Use this to resolve a company NAME into the registration number that lookup_company needs.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Company name or fragment. | |
| limit | No | Max results (default 8). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It discloses that this is a free-text name search scoped to UK Companies House, but it does not describe return format, ordering, pagination, or no-match behavior. It adds some useful context but is not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences. The first gives the core action and scope, the second explains the purpose and relationship to a sibling tool. Every word earns its place, and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter search tool with no output schema, the description is sufficiently complete: it explains the input, the source, and the intended downstream use. It could benefit from mentioning the response shape, but this is not a major gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides full descriptions for both parameters (q and limit) with 100% coverage. The description adds no extra parameter-specific semantics, so it stays at the baseline for schema-covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('search') and resource ('UK Companies House') and defines the output as resolving a company name into the registration number. It also ties directly to the sibling tool 'lookup_company', making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly tells when to use this tool: to resolve a name to a registration number needed by lookup_company. However, it does not explicitly mention when not to use it or how it differs from search_entities, so it lacks full exclusion/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_entitiesAInspect
Search every node in the live WhiteIntel corpus — companies AND people — by name, across all fused sources. This is the lexical search and it always covers the FULL corpus, so it is the fallback whenever semantic_search comes back thin. Returns entity ids you then pass to get_entity or trace_ownership_path. Each hit's source says whether it came from the resolved corpus or a live registry passthrough — it does NOT name the originating registry. For that provenance call get_entity, whose entity.registry_profile names the source register when we hold one — measured 2026-08-11 it was populated on 22 of 32 sampled entities, so expect null sometimes and fall back to linked_records[].registry and connections[].source — or get_dossier, which cites per-record source URLs. Use juris to scope to a country (e.g. gb, ky, us, cy). Reach into the non-UK sources is verified, not assumed: a name search for 'PETROLEO BRASILEIRO' returned FR (siren), BR (lei and br-cnpj) and US (cusip) rows in one response, 2026-08-11.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Entity name or fragment. | |
| risk | No | Optional: filter by risk level. | |
| type | No | Optional: filter by entity kind. | |
| juris | No | Optional: filter by jurisdiction code (e.g. gb, ky, us). | |
| limit | No | Max results (default 20). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that the source field indicates resolved vs live passthrough, does not name the originating registry, suggests get_entity for that, and notes the fallback options. Also provides a concrete example with verified non-UK sources, showing transparency about data coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is overly verbose, containing a long digression about provenance, specific measurement dates, and an example. While the initial sentence is clear, the subsequent details could be streamlined to improve conciseness without losing essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and many siblings, the description covers the primary use case, fallback behavior, result handling, and a key parameter. It could mention potential errors or the exact output format, but it is largely complete for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description only adds extra context for the juris parameter (scoping to a country) and does not elaborate on q, risk, type, or limit. Since the schema already provides descriptions for all parameters (coverage 100%), the added value is limited, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it searches all nodes (companies AND people) by name across all fused sources, and distinguishes it from semantic_search by calling it the fallback lexical search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use it (fallback when semantic_search returns thin results), how to use the returned IDs (pass to get_entity or trace_ownership_path), and mentions the juris parameter for scoping.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
semantic_searchAInspect
Meaning-based entity search over the corpus (BGE-M3 vector ANN over the resolved dossier cards). Finds companies and people whose profile is semantically closest to a natural-language query — a description, a role, a risk pattern — even when no keyword matches. Optional kind (Company/Person/Asset) and jurisdiction (ISO code) filters. Returns entity_id, caption, kind, jurisdiction, risk and a similarity score; feed entity_id into get_dossier / trace_ownership_path. TODAY THIS IS EFFECTIVELY A RISK-LIST SEARCH, NOT A CORPUS SEARCH. The response carries its own coverage object — read it, it is authoritative and it moves. Measured 2026-08-11: embedded 990,055 of a 47,486,969 universe (ratio 0.0208), and per the endpoint's own note that embedded slice is ~99.6% risk-listed and ~97% natural persons. So a query about an ordinary trading company will return sanctioned people and vessels that merely sound related — verified: 'sanctioned russian aluminium holding' returned RU sanctioned SHIPS as its top hits. An empty or off-target result means 'not embedded yet' far more often than 'not found'. ALWAYS pair this with search_entities, which is lexical and covers the full corpus, before concluding anything about an entity's existence. Latency: 6.4s measured on a cold k=5 call — budget for it.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Max hits (default 10). | |
| kind | No | Optional entity-kind filter (Company / Person / Asset / …). | |
| query | Yes | Natural-language search, e.g. 'sanctioned Russian aluminium holding company'. | |
| jurisdiction | No | Optional ISO jurisdiction filter (e.g. GB, RU). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses risk-list bias with measured statistics (embedded 990,055 of 47,486,969, ratio 0.0208; ~99.6% risk-listed, ~97% natural persons), a verified example ('sanctioned russian aluminium holding' returned RU sanctioned ships), the authoritative and moving `coverage` object, and latency (6.4s cold). This is far beyond typical disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: first sentence states purpose, then filters and return fields, followed by caveats, usage guidance, and latency. Each sentence contributes a distinct fact (mechanism, bias, coverage, paired tool, performance). No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema and no annotations, the description fully compensates: it names return fields, the `coverage` object, limitations of the embedded slice, the need to pair with search_entities, and latency. This is enough for an agent to decide whether and how to call the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces kind values (Company/Person/Asset) and jurisdiction as ISO code, and gives a natural-language query example, but it adds little meaning beyond the schema's parameter descriptions. No contradiction or missing parameter context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb and resource: 'Meaning-based entity search over the corpus (BGE-M3 vector ANN over the resolved dossier cards).' It clearly states what it finds (companies and people semantically closest to a natural-language query) and differentiates itself from lexical sibling search_entities. It also names return fields and how to chain results into get_dossier / trace_ownership_path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use this tool versus alternatives: 'ALWAYS pair this with search_entities, which is lexical and covers the full corpus, before concluding anything about an entity's existence.' It also explains that an empty or off-target result means 'not embedded yet' far more often than 'not found', guiding the agent's interpretation and fallback behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_ownership_pathAInspect
Walk the ownership graph upward from a root entity and return the ordered hops connecting it to the ultimate beneficial owner. Use this to answer 'who ultimately controls X?'. Get the root id from search_entities. THE HOP AT THE TOP OF THE LIST IS NOT NECESSARILY THE ULTIMATE OWNER, AND max_depth IS A REQUEST, NOT A PROMISE. Measured 2026-08-11 anonymously: max_depth=6 came back as max_depth: 2, depth_capped: true, plan: 'free' — the walk stopped two hops up and the payload said so only in those two fields. So before you name a UBO, compare hop_count with the RETURNED max_depth and check depth_capped: if the walk was capped and the topmost owner still has owners, you have found an intermediate holder, not the beneficial owner. A paid API key walks deeper. Shape: a single flat hops array (each hop from/fromName/to/toName/role/share/source), not one array per branch. as_observed is a standing caveat: edges carry the date we OBSERVED them in a registry, not a validity period — we hold no ownership end dates, so a link shown here may already have ended.
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | Root entity id to trace from. | |
| max_depth | No | Max hops to request (default 6). The plan lowers it — anonymous callers measured at 2 — so trust the response's `max_depth` / `depth_capped`, not this value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description reveals critical behavior: max_depth is 'a REQUEST, not a PROMISE,' top hop is not necessarily UBO, and includes a measured example of depth capping (max_depth=6 returned as max_depth:2, depth_capped:true). It also discloses the standing caveat on `as_observed` dates, covering data-validity limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries unique information: purpose, use case, root source, critical warnings, response shape, and caveats. It is front-loaded with the main action, but the density makes it slightly less concise than shorter equivalents.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description provides the full response shape (flat `hops` array with fields from/fromName/to/toName/role/share/source), key response fields (`hop_count`, `max_depth`, `depth_capped`), and interpretation guidance. It covers edge cases and data caveats, making it fully contextual.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers both parameters thoroughly (100% coverage), so baseline is 3. The description adds a concrete measured example and the instruction to fetch root from search_entities, enriching semantics but not fundamentally changing the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with 'Walk the ownership graph upward from a root entity and return the ordered hops' – a specific verb+resource+output. It further clarifies the use case with 'who ultimately controls X?' and distinguishes from general graph tools like graph_path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Use this to answer ''who ultimately controls X?''' and instructs 'Get the root id from search_entities.' It also provides detailed guidance on handling depth capping and verifying UBO, effectively telling the agent when to trust the result.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
21 tool updates
v0.7.7- First observed
buy_dossier - First observed
check_offshore_exposure - First observed
claim_dossier - First observed
find_similar - First observed
get_company_details - First observed
get_dossier - First observed
get_entity - First observed
get_financials - First observed
get_payment_link - First observed
get_pricing - First observed
get_pulse - First observed
get_sanctions - First observed
graph_neighbourhood - First observed
graph_path - First observed
lookup_by_identifier - First observed
lookup_company - First observed
resolve - First observed
search_companies - First observed
search_entities - First observed
semantic_search - First observed
trace_ownership_path
TDQS
Most tools have clearly distinct input/output types (e.g., lookup_company vs search_entities vs lookup_by_identifier), and the detailed usage notes remove almost all ambiguity. Minor overlap exists between search_companies and search_entities for UK names, and among graph traversal tools, but descriptions explicitly direct agents to the correct tool.
The majority of tools follow a verb_noun pattern (get_entity, trace_ownership_path, buy_dossier), but a few exceptions—graph_neighbourhood, graph_path, semantic_search, find_similar, resolve—break the pattern. These are minor deviations in an otherwise consistent naming scheme.
21 tools is on the higher end, but each tool serves a distinct function across entity lookup, graph analysis, sanctions screening, financials, and purchasing. The payment-related tools (pricing, buy, link, claim) add count but form a logical workflow. The number is appropriate for the broad domain, though slightly heavier than typical.
The toolset covers the full range of entity intelligence: multiple search/resolution methods, deep dives (dossier, details, financials), graph analysis (paths, neighborhoods, UBO tracing), sanctions and offshore checks, activity feed, similarity search, and a complete purchase flow. No critical operation appears missing for the stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Global B2B intelligence for AI agents: 35M+ companies, 1.6M sanctions, KYB pack. 78 tools.
US public-records intelligence for AI agents — companies, SEC, courts, spending, licenses.
Sanctions screening, KYB, identifier validation, jurisdiction risk & secret scanning for AI agents
Entity verification, sanctions screening, and trust scoring for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceProvides real-time company verification and corporate intelligence by accessing global registries like UK Companies House, Singapore ACRA, and OpenCorporates. It enables AI agents to perform KYC tasks, retrieve company profiles, and conduct automated risk assessments for due diligence workflows.249MIT
- AlicenseAqualityCmaintenanceStructured business intelligence for AI agents. 5.5M verified entities across 34 countries, 40.3M BORME mercantile acts, EU VAT validation, GLEIF, healthcare registries. 20 tools.61MIT
- AlicenseAqualityDmaintenanceUnmodified government company data from 27 registries, live. Cross-border UBO chain walker for AI agents. 60+ tools covering GB, IE, NO, FR, DE, NL, PL, BE, CH, LI, MC, IM, IS, CY, AU, NZ, CA, TW, HK, MY, FI, CZ, ES, IT, KR, US — raw upstream fields preserved, no LLM extraction.1017Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to access official European business data across 15 EU countries, including company lookups, VAT validation, sanctions screening, and KYB reports.11,5021MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Hei33enberg/WhiteIntel-OS'
If you have feedback or need assistance with the MCP directory API, please join our Discord server