Skip to main content
Glama

campaignstack_list_voice_experiment_results

Read-onlyIdempotent

Reply-rate numbers for the voice vs generic experiment: sends, replies, acceptance counts, reply rate with a 95% confidence interval per arm, the relative lift, and whether the lift thesis is validated at volume. Sliceable by channel, craft kind, and profile version.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
channelNoSlice by channel, e.g. linkedin
craftKindNoSlice by craft kind: note, message, comment, reply
workspaceIdNoDefaults to the API key's workspace
profileVersionNoSlice by voice profile version

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. First observed

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is established. The description adds useful context about the output content (confidence interval, lift, validation) and sliceability, but it does not disclose behavioral details like default time windows, empty-result behavior, or pagination. This adds some value beyond annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler; the core purpose and output metrics are front-loaded in the first sentence, with filtering options in the second. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, no-required-params metrics listing, the description covers the returned values and all relevant slicing dimensions; the schema covers the parameters. The lack of an output schema is compensated by the explicit metric list, though details like the time range or the definition of 'at volume' are not specified, leaving slightly more to infer than ideal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters have schema descriptions with 100% coverage, so the schema does the heavy lifting. The description's 'Sliceable by channel, craft kind, and profile version' largely restates the schema fields, and it adds no format, enumeration, or default semantics beyond the schema's workspaceId note. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (voice vs generic experiment) and enumerates exactly what is returned: sends, replies, acceptance counts, reply rate with 95% CI per arm, relative lift, and validation status. This clearly distinguishes it from the many other list_* siblings, which concern leads, campaigns, accounts, etc. The only minor weakness is that it starts with a noun phrase rather than an explicit verb, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when an agent needs reply-rate and lift metrics for the voice vs generic experiment, and it states the available slicing dimensions. However, it offers no explicit when-to-use vs alternatives or any exclusions (e.g., when to prefer get_campaign_metrics or list_split_optimization_logs). This is more than no guidance but less than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.6/5.0
Disambiguation3/5

Many tools share the same verb prefix (create_, list_, update_, get_) across closely related resources, so pairs like add_lead_to_external_list vs add_lead_to_sequence, create_signal_agent vs create_signal_watch, and approve_review vs approve_content_post can be confused. The descriptions are unusually detailed and cross-referenced, which mitigates but does not eliminate the ambiguity inherent in a 282-tool surface.

Naming Consistency4/5

Virtually every tool follows the campaignstack_verb_noun snake_case pattern, which is highly predictable. Minor deviations exist: destructive operations mix remove_ and delete_ (remove_lead_list vs delete_campaign), AI generation uses both craft_ and generate_, and the seo_/search_console_ subdomains introduce a second prefix convention.

Tool Count1/5

282 tools is an extreme mismatch by any reasonable standard, exceeding the 50+ threshold by more than 5x. Even for a full B2B outreach platform, this surface is far too large and would be better consolidated into higher-level operations or grouped sub-servers.

Completeness4/5

The tool surface is impressively comprehensive, covering campaigns, workflows, leads, content, ads, SEO, integrations, billing, and more with CRUD-level depth. Minor gaps remain: no single-ICP getter, no direct pause/delete for search watches, and no explicit delete for ad campaigns (only archive via update).

Resources