TowerWatch Ops Agent MCP Server
Provides the ability to run network speed tests via Speedtest.net.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@TowerWatch Ops Agent MCP Serverquery metrics for latency in last hour"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TowerWatch Ops Agent
An agent layer over TowerWatch ā the network-quality monitoring project ā built to demonstrate the three capabilities an enterprise agent-engineering loop needs: evaluation suites, cost/latency-aware model choice, and tool retrieval. One repo, one coherent story:
"I took my public monitoring project and built the agent layer an enterprise would need around it: an instrumented MCP server with defined SLIs, an eval harness in CI that catches seeded regressions, a cost-aware model router, and semantic tool retrieval with measured selection precision."
At a glance
Domain: TowerWatch's network-monitoring data, exposed as agent tools.
Transport: stdio first; stateless streamable HTTP as a stretch goal.
Observability: OpenTelemetry from the first tool call, into a Prometheus/Grafana stack.
Tool surface: seven tools ā
query_metrics,analyze_window,compare,query_log_events,get_monitor_status,get_runbook,run_speedtest. Contracts indocs/design/.Status: š” Phase 1 in progress ā the server runs and one of seven tools is built; no Phase 1 acceptance criterion is met yet. See Status.
Why this project
It fills the gap between "I read about agent evaluation and routing" and "I built and measured it." Every artifact ā eval tables, benchmark numbers, precision@k charts ā is a number personally collected, not a claim from study. The domain is real data from a project the author already owns, so the story is "I extended my own production-style system," not "I did a tutorial."
The build runs under the author's own Agent Collaboration Principles: every phase's definition-of-done is a set of independently checkable artifacts ā a command that runs, a file that exists, a dashboard that renders. No "trust me, it works."
Related MCP server: production-grade-mcp-agentic-system
The three phases
The project is one build in three strictly-sequenced phases. Full specs live in
docs/specs/; the build plan is the index.
Requirements were defined upfront in a planning process and built to as a contract ā
the specs came first, the tool contracts were derived from them, and the
ADRs record every decision that shaped the surface.
Phase | Ships | Spec |
1 | Instrumented MCP server over TowerWatch data + defined SLIs + cross-model cost/latency bench | |
2 | Golden-set + rubric eval harness in CI that catches a seeded regression | |
3 | Cost-aware model router + semantic tool retrieval with measured selection precision | |
Cross-cutting | Agent-facing docs, in-repo skills, ADRs, and a measured onboarding eval ā incremental alongside the phases, never blocking |
Sequence is strict: Phase 2's evals score Phase 3's router. Don't reorder. The cross-cutting layer is the exception ā it lands incrementally and gates nothing.
Repository layout
towerwatch-ops-agent/
āāā README.md # this file ā human-facing
āāā CLAUDE.md # agent-facing anchor (read first if you're an agent)
āāā pyproject.toml # PEP 621 single source of truth ā deps, tooling config
āāā docs/
ā āāā architecture.md # intended shape (stub ā not built yet)
ā āāā specs/ # the governing build plan + 4 requirement specs
ā āāā design/ # locked tool contracts (00ā11) ā authoritative
ā āāā adr/ # architecture decision records
ā āāā production-path.md # personal-scale choices vs. enterprise needs
āāā src/towerwatch_ops_agent/ # server, config, domain/, tools/, telemetry/
āāā tests/ # pytest suite ā 95 tests
āāā fixtures/stub/ # hand-authored stub corpus (not the real one)
āāā RATIONALE.md # deliberate choices that read as defectsQuick start
The server runs and serves
query_metrics. The other six tools are not built yet.
# From repo root. uv manages the environment and lockfile.
uv sync # create .venv, install deps from pyproject.toml
uv run python -m towerwatch_ops_agent # (Phase 1) launch the MCP server over stdioTesting the server interactively (Phase 1) uses the MCP Inspector:
npx @modelcontextprotocol/inspector uv run python -m towerwatch_ops_agentStatus
š” Phase 1 in progress. The MCP server runs over stdio and serves
query_metrics end to end against a fixture. None of Phase 1's five acceptance
criteria are met yet ā see spec-phase1-mcp-server.md
for the gate list.
Built and running:
Directory skeleton,
pyproject.toml,.gitignore, MIT licenseREADME,
CLAUDE.md(with binding invariants), architecture stubThe build plan and all four requirement specs in
docs/specs/Locked tool contracts ā
docs/design/00ā11: conventions, seven tool docs, skills interfaces, span schema, fixture manifest, eval designADRs ā
docs/adr/, the decisions behind the tool surfaceMCP server + composition root ā
server.py,config.py, stdio transportquery_metricsā 1 of 7 tools, with thedata_statusenvelope enforcedFixtureClient+ manifest loader ā ADR-0002's dual-mode seam, fixture side onlySpan instrumentation ā one span per tool call, secrets structurally excluded
CI workflow ā ruff, format, pyright, pytest on every PR branch head
RATIONALE.mdā deliberate choices a reviewer would otherwise report as defects
Deferred (not yet built ā see CLAUDE.md for the phase gates):
Six remaining tools ā
analyze_window,compare,query_log_events,get_monitor_status,get_runbook,run_speedtestGrafanaCloudClientā the live half of theDataClientProtocolCurated fixture corpus ā
fixtures/stub/is a two-window hand-authored stub proving the format only, not the real deterministic corpusOTel exporter + SLI dashboard ā spans are emitted but go nowhere; no
MeterProvider, so no duration histogramsdef_tokens.mdā the tool-def token budget measurement (script exists, never run)bench.mdā cross-model cost/latency benchPhase 2 ā eval harness + CI + seeded-regression showpiece
Phase 3 ā model router + semantic tool retrieval
In-repo skills under
.claude/skills/ādiagnose-rca,evidence-pack, plus the golden-path skills (add-tool,run-evals) created when first walked manuallyMeasured onboarding eval (
docs/onboarding-eval.md) ā first run after Phase 1
For AI assistants
If you're an agent working in this repo, read CLAUDE.md first. It
carries the phase sequence, the stateless-gates working standard, and an explicit map of
what exists versus what is still a stub, so you don't reason about code that isn't there
yet. RATIONALE.md records the deliberate choices that read as defects on sight ā read it
before reporting one.
Available Tools
1 tooltowerwatch_query_metricsARead-only
Raw time-series data points from TowerWatch network monitoring.
Pick this when you need the actual numbers ā specific values, series, timestamps ā and you will do your own reasoning over them. If you want a judgment about a window (is it degraded, and against what reference), use analyze_window instead.
Returns downsampled [timestamp, value] pairs per metric, plus data_status. Read data_status before the numbers: 'empty_window' means collected here with nothing in range (a true negative), while 'not_collected' means this site never collects it ā no evidence, so do not infer that anything is healthy.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present only when data_status is 'error'. |
| series | No | Metric name to its downsampled points. Empty unless data_status is ok. |
| truncated | No | True when more points exist beyond this page. |
| data_status | Yes | ok=data present; empty_window=collected here, none in range (true negative); not_collected=site never collects this (NO evidence ā do not infer health); partial=some groups missing; error=see message. |
| coverage_notes | No | Why data is missing or partial, in plain language. |
| next_page_token | No | Pass back as page_token to continue. Null when complete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and destructiveHint. The description adds meaningful behavioral context by explaining data_status semantics: 'empty_window' as a true negative versus 'not_collected' as no evidence, which is critical for interpreting results. It also discloses downsampling behavior and per-series output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, usage selection, return format, and an important caveat about data_status. The structure is front-loaded and the caveat is placed where it will be read before acting on numbers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool, the description covers when to use it, what it returns, and the crucial data_status interpretation. Pagination and request shape are documented in the schema, and there is an output schema, so the description is complete enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the primary request parameters such as site, start, end, metric_group, or pagination. It only implies per-metric and downsampled behavior. The nested schema helps, but the description itself does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it returns raw time-series data points as downsampled [timestamp, value] pairs per metric, and explicitly distinguishes itself from analyze_window by saying this tool is for actual numbers while the sibling is for judgments. This gives an agent a clear, specific understanding of the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to pick this tool when actual numbers are needed and the agent will do its own reasoning, and directs users to analyze_window when they want a judgment about a window. This is clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v0.0.0- First observed
towerwatch_query_metrics
TDQS
With only one tool defined, there is no possibility of confusion between overlapping tools. The tool's purpose is clearly described, though it references a missing 'analyze_window' tool that does not exist in the server.
A single tool name following a clear prefix+verb_noun pattern (towerwatch_query_metrics) provides no inconsistency issues. There is no mix of conventions to evaluate.
A server with only one tool is very thin for a monitoring domain, especially since the description explicitly references a second tool ('analyze_window') that is absent. The scope is too narrow for an agent to perform useful monitoring workflows.
The tool only returns raw time series data and explicitly defers judgment to 'analyze_window', which is not implemented. This is a significant gap: agents cannot obtain window-level health assessments, and the missing referenced tool creates a dead end.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for building and testing AI agents with multi-model experimentation and insights.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
471Monitoring for the agent economy ā liveness, latency, trust scoring for MCP endpoints
1MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP-native agent evaluation and observability server. Log traces, evaluate output quality with 12 built-in rules (PII detection, prompt injection, cost thresholds), and track agent costs. Real-time dashboard, OTel-compatible spans. Self-hosted, MIT licensed.91299MIT
- AlicenseNot gradedqualityDmaintenanceA production-grade MCP server designed for multi-tenant, authenticated, and observable AI agent systems, enabling secure tool execution across heterogeneous data sources.62MIT
- AlicenseAqualityBmaintenanceAn MCP server exposing 72 tools across 26 homelab services, enabling LLMs to monitor and manage infrastructure, media, storage, and networking with a single endpoint.16MIT
- AlicenseAqualityDmaintenanceAn MCP server that exposes live network monitoring data as Resources and diagnostic capabilities as Tools, letting AI assistants query network health conversationally.6MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/kemosabe102/towerwatch-ops-agent'
If you have feedback or need assistance with the MCP directory API, please join our Discord server