CtrlTest MCP Server
The CtrlTest MCP Server provides automated control system evaluation and regression testing for PID and bio-inspired controllers in flapping-wing aircraft systems, exposing a standardized API (ctrltest.analyze_pid) for comprehensive performance analysis.
Core Capabilities:
PID Controller Analysis - Evaluate PID gains against plant dynamics (natural frequency, damping ratio) to measure overshoot, settling time, integral squared error (ISE), and Lyapunov stability margins
High-Fidelity Simulation Integration - Ingest data from diffSPH (differentiable smoothed-particle hydrodynamics) and Foam-Agent CFD simulations via
extra_metricsfor realistic performance assessmentMulti-Modal Scoring - Fuse analytic models with high-fidelity simulation data when both diffSPH and Foam-Agent metrics are available
Gust Rejection Analysis - Quantify disturbance handling through gust detection latency, bandwidth, and rejection percentage metrics
Energy Efficiency Assessment - Measure CPG (Central Pattern Generator) baseline vs. consumed energy and calculate energy reduction percentages
Mix-of-Experts Evaluation - Compute switching penalties, latency budgets, and energy consumption for adaptive control architectures
Automated Integration - Support continuous performance monitoring, automated gain tuning, and controller optimization through STDIO/HTTP API integration with ToolHive, RL algorithms, and evolutionary systems
Structured Output - Export JSON-formatted results with provenance metadata for dashboards, logging, and downstream analysis
Provides a REST API interface for the control system evaluation service, allowing HTTP-based access to controller benchmarking and scoring functionality.
Enables visualization of controller performance metrics and analytics data exported from PID evaluations and comparative controller assessments.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CtrlTest MCP Serverevaluate PID gains kp=2.0, ki=0.5, kd=0.12 for a second-order plant with natural frequency 3.2Hz"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ctrltest-mcp - Flight-control regression lab for MCP agents
TL;DR: Evaluate PID and bio-inspired controllers against analytic or diffSPH/Foam-Agent data through MCP, logging overshoot, energy, and gust metrics automatically.
Table of contents
Related MCP server: shewhart-mcp
What it provides
Scenario | Value |
Analytic PID benchmarking | Run closed-form plant models and produce overshoot/settling/energy metrics without manual scripting. |
High-fidelity scoring | Ingest logged data from Foam-Agent or diffSPH runs and fuse it into controller evaluations. |
MCP integration | Expose the scoring API via STDIO/HTTP so ToolHive or other clients can automate gain tuning and generate continuous performance scorecards. |
Quickstart
uv pip install "git+https://github.com/Three-Little-Birds/ctrltest-mcp.git"Run a PID evaluation:
from ctrltest_mcp import (
ControlAnalysisInput,
ControlPlant,
ControlSimulation,
PIDGains,
evaluate_control,
)
request = ControlAnalysisInput(
plant=ControlPlant(natural_frequency_hz=3.2, damping_ratio=0.35),
gains=PIDGains(kp=2.0, ki=0.5, kd=0.12),
simulation=ControlSimulation(duration_s=3.0, sample_points=400),
setpoint=0.2,
)
response = evaluate_control(request)
print(response.model_dump())Typical outputs (analytic only):
{
"overshoot": -0.034024863556091134,
"ise": 0.008612387509182674,
"settling_time": 3.0,
"gust_detection_latency_ms": 0.8,
"gust_detection_bandwidth_hz": 1200.0,
"gust_rejection_pct": 0.396,
"cpg_energy_baseline_j": 12.0,
"cpg_energy_consumed_j": 7.8,
"cpg_energy_reduction_pct": 0.35,
"lyapunov_margin": 0.12,
"moe_switch_penalty": 0.135,
"moe_latency_ms": 12.72,
"moe_energy_j": 3.9,
"multi_modal_score": null,
"extra_metrics": null,
"metadata": {"solver": "analytic"}
}The analytic plant example above clips
settling_timeat the requested simulation duration (duration_s=3.0). Increase the horizon if you need the loop to settle fully before computing that metric.
Run as a service
CLI (STDIO transport)
uvx ctrltest-mcp # runs the MCP over stdio
# or just python -m ctrltest_mcpUse python -m ctrltest_mcp --describe to print basic metadata without starting the server.
FastAPI (REST)
uv run uvicorn ctrltest_mcp.fastapi_app:create_app --factory --port 8005python-sdk tool (STDIO / MCP)
from mcp.server.fastmcp import FastMCP
from ctrltest_mcp.tool import build_tool
mcp = FastMCP("ctrltest-mcp", "Flapping-wing control regression")
build_tool(mcp)
if __name__ == "__main__":
mcp.run()ToolHive smoke test
Run the integration script from your workspace root:
uvx --with 'mcp==1.20.0' python scripts/integration/run_ctrltest.pyThe smoke test runs the analytic path by default. To exercise high-fidelity scoring, stage Foam-Agent archives under logs/foam_agent/ and diffSPH gradients under logs/diffsph/ before launching the script.
Agent playbook
Gust rejection - feed archived diffSPH gradients (
diffsph_metrics) and Foam-Agent archives (paths returned by those services) to quantify adaptive CPG improvements.Controller comparison - log analytics for multiple PID gains, export JSONL evidence, and visualise in Grafana.
Policy evaluation - integrate with RL or evolutionary algorithms; metrics are structured for automated scoring.
Stretch ideas
Extend the adapter for PteraControls (planned once upstream Python bindings are published).
Drive the MCP from
scripts/fitnessto populate nightly scorecards.Combine with
migration-mcpto explore route-specific disturbance budgets.
Accessibility & upkeep
Run
uv run pytest(tests mock diffSPH/Foam-Agent inputs and assert deterministic analytic results).Keep metric schema changes documented—downstream dashboards rely on them.
Metric schema at a glance
Field | Units | Notes |
| radians | peak response minus setpoint |
| rad²·s | integral squared error |
| seconds | first time error stays within tolerance |
| milliseconds | detector latency |
| hertz | detector bandwidth |
| 0–1 | fraction of disturbance rejected |
| joules | energy pre-adaptation |
| joules | energy post-adaptation |
| 0–1 | energy reduction ratio |
| unitless | stability margin |
| unitless | cost weight × switches |
| milliseconds | latency budget after switching |
| joules | mix-of-experts energy draw |
| unitless | only when both diffSPH & Foam metrics are present |
| varies | raw diffSPH/Foam metrics merged in |
Example of fused high-fidelity metrics:
{
"extra_metrics": {
"force_gradient_norm": 0.87,
"lift_drag_ratio": 18.4
},
"multi_modal_score": 0.047,
"metadata": {"solver": "analytic"}
}Contributing
uv pip install --system -e .[dev]uv run ruff check .anduv run pytestShare sample metrics in PRs so reviewers can sanity-check improvements quickly.
MIT license - see LICENSE.
Available Tools
1 toolctrltest.analyze_pidC
Score PID gains for a flapping-wing plant. Provide plant dynamics and optional gradients/metadata. Returns key control metrics plus provenance. Example input: {"plant":{"natural_frequency_hz":4.2,"damping_ratio":0.45},"gains":{"kp":1.1,"ki":0.2,"kd":0.08},"diffsph_metrics":{"force_gradient_norm":1.9}}
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| ise | Yes | |
| metadata | No | |
| overshoot | Yes | |
| moe_energy_j | Yes | |
| extra_metrics | No | |
| settling_time | Yes | |
| moe_latency_ms | Yes | |
| lyapunov_margin | Yes | |
| multi_modal_score | No | |
| gust_rejection_pct | Yes | |
| moe_switch_penalty | Yes | |
| cpg_energy_baseline_j | Yes | |
| cpg_energy_consumed_j | Yes | |
| cpg_energy_reduction_pct | Yes | |
| gust_detection_latency_ms | Yes | |
| gust_detection_bandwidth_hz | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions what the tool returns ('key control metrics plus provenance'), it doesn't describe important behavioral aspects like computational requirements, accuracy limitations, whether it's a simulation or real-time analysis, error conditions, or performance characteristics. The description is insufficient for a complex control analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise with two sentences plus an example. The first sentence states the purpose and required inputs, the second describes the output, and the example provides concrete illustration. However, the example could be more focused on illustrating the structure rather than specific values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complex parameter structure (1 top-level parameter with 10 nested properties across multiple objects), 0% schema description coverage, and no annotations, the description is incomplete. While an output schema exists (which helps), the description doesn't adequately explain the sophisticated control engineering concepts involved or the tool's operational context. The example helps but doesn't compensate for the missing conceptual explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'plant dynamics and optional gradients/metadata' which aligns with the input schema's structure. However, with 0% schema description coverage, the description doesn't adequately explain the complex nested parameter structure (plant, gains, simulation, setpoint, gust_detector, adaptive_cpg, moe_router, diffsph_metrics, foam_metrics, prefer_high_fidelity). The example input shows some parameters but doesn't cover the full complexity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Score PID gains for a flapping-wing plant' with specific verb ('Score') and resource ('PID gains'). It mentions providing 'plant dynamics and optional gradients/metadata' and returning 'key control metrics plus provenance'. However, without sibling tools, we cannot assess differentiation from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions the tool's function but offers no context about prerequisites, typical use cases, or limitations. The example input shows what data to provide, but doesn't explain when this analysis would be appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v1.0.0- Changed
ctrltest.analyze_pid2 fields changed- added
Input schema / $defsAdded value: +{ + "AdaptiveCPGConfig": { + "properties": { + "energy_baseline_j": { + "default": 12, + "exclusiveMinimum": 0, + "title": "Energy Baseline J", + "type": "number" + }, + "energy_reduction_pct": { + "default": 0.35, + "maximum": 0.95, + "minimum": 0, + "title": "Energy Reduction Pct", + "type": "number" + }, + "lyapunov_margin": { + "default": 0.12, + "minimum": 0, + "title": "Lyapunov Margin", + "type": "number" + }, + "target_rejection_pct": { + "default": 0.45, + "maximum": 0.95, + "minimum": 0, + "title": "Target Rejection Pct", + "type": "number" + } + }, + "title": "AdaptiveCPGConfig", + "type": "object" + }, + "ControlAnalysisInput": { + "properties": { + "adaptive_cpg": { + "$ref": "#/$defs/AdaptiveCPGConfig" + }, + "diffsph_metrics": { + "anyOf": [ + { + "additionalProperties": true, + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Diffsph Metrics" + }, + "foam_metrics": { + "anyOf": [ + { + "additionalProperties": true, + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Foam Metrics" + }, + "gains": { + "$ref": "#/$defs/PIDGains" + }, + "gust_detector": { + "$ref": "#/$defs/GustDetectorConfig" + }, + "moe_router": { + "$ref": "#/$defs/MoERouterConfig" + }, + "plant": { + "$ref": "#/$defs/ControlPlant" + }, + "prefer_high_fidelity": { + "default": true, + "description": "Attempt to use PteraControls when available before falling back to the analytic surrogate.", + "title": "Prefer High Fidelity", + "type": "boolean" + }, + "setpoint": { + "default": 0, + "title": "Setpoint", + "type": "number" + }, + "simulation": { + "$ref": "#/$defs/ControlSimulation" + } + }, + "required": [ + "plant", + "gains" + ], + "title": "ControlAnalysisInput", + "type": "object" + }, + "ControlPlant": { + "properties": { + "damping_ratio": { + "minimum": 0, + "title": "Damping Ratio", + "type": "number" + }, + "natural_frequency_hz": { + "exclusiveMinimum": 0, + "title": "Natural Frequency Hz", + "type": "number" + }, + "settling_tolerance_rad": { + "default": 0.02, + "exclusiveMinimum": 0, + "title": "Settling Tolerance Rad", + "type": "number" + }, + "trim_setpoint": { + "default": 0, + "title": "Trim Setpoint", + "type": "number" + } + }, + "required": [ + "natural_frequency_hz", + "damping_ratio" + ], + "title": "ControlPlant", + "type": "object" + }, + "ControlSimulation": { + "properties": { + "duration_s": { + "default": 5, + "exclusiveMinimum": 0, + "title": "Duration S", + "type": "number" + }, + "sample_points": { + "default": 500, + "maximum": 5000, + "minimum": 50, + "title": "Sample Points", + "type": "integer" + } + }, + "title": "ControlSimulation", + "type": "object" + }, + "GustDetectorConfig": { + "properties": { + "bandwidth_hz": { + "default": 1200, + "minimum": 10, + "title": "Bandwidth Hz", + "type": "number" + }, + "latency_ms": { + "default": 0.8, + "minimum": 0.1, + "title": "Latency Ms", + "type": "number" + }, + "sensitivity": { + "default": 0.88, + "maximum": 1, + "minimum": 0, + "title": "Sensitivity", + "type": "number" + } + }, + "title": "GustDetectorConfig", + "type": "object" + }, + "MoERouterConfig": { + "properties": { + "energy_budget_j": { + "default": 6, + "exclusiveMinimum": 0, + "title": "Energy Budget J", + "type": "number" + }, + "latency_budget_ms": { + "default": 12, + "exclusiveMinimum": 0, + "title": "Latency Budget Ms", + "type": "number" + }, + "switch_cost_weight": { + "default": 0.045, + "minimum": 0, + "title": "Switch Cost Weight", + "type": "number" + }, + "switch_events": { + "default": 3, + "minimum": 0, + "title": "Switch Events", + "type": "integer" + } + }, + "title": "MoERouterConfig", + "type": "object" + }, + "PIDGains": { + "properties": { + "kd": { + "default": 0, + "minimum": 0, + "title": "Kd", + "type": "number" + }, + "ki": { + "default": 0, + "minimum": 0, + "title": "Ki", + "type": "number" + }, + "kp": { + "minimum": 0, + "title": "Kp", + "type": "number" + } + }, + "required": [ + "kp" + ], + "title": "PIDGains", + "type": "object" + } +} - added
Input schema / titleAdded value: +"analyzeArguments"
1 tool update
- First observed
ctrltest.analyze_pid
TDQS
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as analyzing PID gains for a specific type of plant, making it distinct by default.
The single tool name follows a consistent pattern with a clear namespace prefix (ctrltest) and descriptive action (analyze_pid). With only one tool, naming consistency is inherently perfect as there are no other tools to compare against.
A single tool is generally too few for most server purposes, as it limits functionality and forces agents to rely on one operation. While it might be appropriate for a highly specialized task, the server's name 'CtrlTest MCP Server' suggests a broader control testing domain where more tools would be expected for completeness.
The server appears focused on control system testing, but with only one analysis tool, there are significant gaps. Missing operations likely include tools for setting up tests, running simulations, comparing results, or managing test configurations, making the surface severely incomplete for the implied domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Deterministic engineering solvers: biquad filter design, room eigenmodes, LLM VRAM fit.
51Evidence-backed architecture-quality analysis for Python agent applications.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
3rd Generation Testing (3TG) — generate deterministic test suites from Markdown spec tables via MCP.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA dual-track testing server that combines CLI test execution with Playwright-based browser testing and persistent SQLite logging. It enables automated test pipelines, Git integration, and evidence-based requirement generation to streamline the development lifecycle.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to perform statistical process control calculations using validated, deterministic tools such as control charts, capability analysis, and tolerance intervals.3MIT
- AlicenseNot gradedqualityAmaintenanceValidated PNT-resilience simulator over MCP — SGP4/SDP4 orbit propagation, IAU reference frames, GNSS availability/DOP, GNSS/INS fusion, ARAIM integrity, and Allan deviations, with results validated against AIAA/IGS/SOFA/NIST reference data.6AGPL 3.0
- AlicenseBqualityBmaintenanceAn MCP-style stdio server for evaluating AI agent outputs, enabling CI-friendly quality gates, regression comparisons, and canary promotion decisions.3MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Three-Little-Birds/ctrltest-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server