Skip to main content
Glama

Evaluate Model compute

model_evaluate

Read-only computeModel. Does not write MySQL. View ids allowed. Default period is calendar start (not today). periodId and range* are mutually exclusive. Cell cap 2000.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNo
setIdNo
sliceNo
periodIdNo
rangeEndNo
rangeStartNo
includeEmptyNo
parameterIdsNo
includeInputsNo

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Added

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers meaningful behavioral context: read-only semantics, no MySQL writes, permission scope, default period behavior, mutual exclusivity, and a cell cap of 2000. This goes well beyond a bare restatement of purpose, though it still omits details about return structure or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: read-only safety, write scope, permissions, default behavior, a constraint, and a boundary limit. The description is compact, front-loaded with the most important trait, and contains no filler or redundant explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high parameter count, nested slice object, absence of annotations, and lack of an output schema, the description is not fully complete. It captures key constraints and defaults but leaves many parameter meanings and the return shape undocumented, so an agent may still need to infer important calling details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds valuable semantics for periodId and range* by noting mutual exclusivity and the default period, but it does not clarify id vs setId, slice, includeEmpty, parameterIds, or includeInputs. With 9 parameters, this leaves substantial ambiguity unresolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a read-only compute operation on a model, immediately distinguishing it from sibling mutation tools like model_upsert_parameter. The phrase 'Does not write MySQL' reinforces the non-mutating scope. It does not fully explain the term 'computeModel' but the intent is unambiguous enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful usage constraints such as 'View ids allowed', 'Default period is calendar start (not today)', and 'periodId and range* are mutually exclusive'. However, it does not explicitly state when to use this tool versus alternatives or when not to use it, leaving much of the routing decision to implication.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.2/5.0
Disambiguation2/5

Many tools are mirrored across `artifacts_*` and `roadflow_*` with near-identical names and behavior, and within each family `get`, `export_json`, and `export_markup` overlap in what they return. Descriptions identify the target workspace, but an agent must carefully inspect prefixes and formats to avoid misselection.

Naming Consistency4/5

The set consistently uses lowercase snake_case with a domain prefix and predictable verbs like get, list, create, open, export, and apply. Minor deviations are bare commands (`new`, `discard`, `status`) and the parallel `artifacts_*`/`roadflow_*` prefixes, which make names look duplicated.

Tool Count3/5

At 23 tools, the surface lands in the heavy 16-25 range and feels padded because many operations are duplicated for two workspace types. Each subsystem alone would have a reasonable count, but combined the set is bloated.

Completeness4/5

The toolset covers the full workspace lifecycle: create, read, update via apply, discard, share, version, status, and multiple export formats. Missing cloud deletion and fine-grained element editing are minor gaps that can be worked around with full-state apply/export.

Resources