Skip to main content
Glama

Test Runner MCP

A Model Context Protocol (MCP) server for running and parsing test results from multiple testing frameworks. This server provides a unified interface for executing tests and processing their outputs, supporting:

  • Bats (Bash Automated Testing System)

  • Pytest (Python Testing Framework)

  • Flutter Tests

  • Jest (JavaScript Testing Framework)

  • Go Tests

  • Rust Tests (Cargo test)

  • Generic (for arbitrary command execution)

Installation

npm install test-runner-mcp

Related MCP server: Tailscale MCP Server

Prerequisites

The following test frameworks need to be installed for their respective test types:

Usage

Configuration

Add the test-runner to your MCP settings (e.g., in claude_desktop_config.json or cline_mcp_settings.json):

{
  "mcpServers": {
    "test-runner": {
      "command": "node",
      "args": ["/path/to/test-runner-mcp/build/index.js"],
      "env": {
        "NODE_PATH": "/path/to/test-runner-mcp/node_modules",
        // Flutter-specific environment (required for Flutter tests)
        "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
        "PUB_CACHE": "/Users/username/.pub-cache",
        "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
      }
    }
  }
}

Note: For Flutter tests, ensure you replace:

  • /opt/homebrew/Caskroom/flutter/3.27.2/flutter with your actual Flutter installation path

  • /Users/username/.pub-cache with your actual pub cache path

  • Update PATH to include your system's actual paths

You can find these values by running:

# Get Flutter root
flutter --version

# Get pub cache path
echo $PUB_CACHE   # or default to $HOME/.pub-cache

# Get Flutter binary path
which flutter

Running Tests

Use the run_tests tool with the following parameters:

{
  "command": "test command to execute",
  "workingDir": "working directory for test execution",
  "framework": "bats|pytest|flutter|jest|go|rust|generic",
  "outputDir": "directory for test results",
  "timeout": "test execution timeout in milliseconds (default: 300000)",
  "env": "optional environment variables",
  "securityOptions": "optional security options for command execution"
}

Example for each framework:

// Bats
{
  "command": "bats test/*.bats",
  "workingDir": "/path/to/project",
  "framework": "bats",
  "outputDir": "test_reports"
}

// Pytest
{
  "command": "pytest test_file.py -v",
  "workingDir": "/path/to/project",
  "framework": "pytest",
  "outputDir": "test_reports"
}

// Flutter
{
  "command": "flutter test test/widget_test.dart",
  "workingDir": "/path/to/project",
  "framework": "flutter",
  "outputDir": "test_reports",
  "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
  "PUB_CACHE": "/Users/username/.pub-cache",
  "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
}

// Jest
{
  "command": "jest test/*.test.js",
  "workingDir": "/path/to/project",
  "framework": "jest",
  "outputDir": "test_reports"
}

// Go
{
  "command": "go test ./...",
  "workingDir": "/path/to/project",
  "framework": "go",
  "outputDir": "test_reports"
}

// Rust
{
  "command": "cargo test",
  "workingDir": "/path/to/project",
  "framework": "rust",
  "outputDir": "test_reports"
}

// Generic (for arbitrary commands, CI/CD tools, etc.)
{
  "command": "act -j build",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports"
}

// Generic with security overrides
{
  "command": "sudo docker-compose -f docker-compose.test.yml up",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports",
  "securityOptions": {
    "allowSudo": true
  }
}

Security Features

The test-runner includes built-in security features to prevent execution of potentially harmful commands, particularly for the generic framework:

  1. Command Validation

    • Blocks sudo and su by default

    • Prevents dangerous commands like rm -rf /

    • Blocks file system write operations outside of safe locations

  2. Environment Variable Sanitization

    • Filters out potentially dangerous environment variables

    • Prevents overriding critical system variables

    • Ensures safe path handling

  3. Configurable Security

    • Override security restrictions when necessary via securityOptions

    • Fine-grained control over security features

    • Default safe settings for standard test usage

Security options you can configure:

{
  "securityOptions": {
    "allowSudo": false,        // Allow sudo commands
    "allowSu": false,          // Allow su commands
    "allowShellExpansion": true, // Allow shell expansion like $() or backticks
    "allowPipeToFile": false   // Allow pipe to file operations (> or >>)
  }
}

Flutter Test Support

The test runner includes enhanced support for Flutter tests:

  1. Environment Setup

    • Automatic Flutter environment configuration

    • PATH and PUB_CACHE setup

    • Flutter installation verification

  2. Error Handling

    • Stack trace collection

    • Assertion error handling

    • Exception capture

    • Test failure detection

  3. Output Processing

    • Complete test output capture

    • Stack trace preservation

    • Detailed error reporting

    • Raw output preservation

Rust Test Support

The test runner provides specific support for Rust's cargo test:

  1. Environment Setup

    • Automatically sets RUST_BACKTRACE=1 for better error messages

  2. Output Parsing

    • Parses individual test results

    • Captures detailed error messages for failed tests

    • Identifies ignored tests

    • Extracts summary information

Generic Test Support

For CI/CD pipelines, GitHub Actions via act, or any other command execution, the generic framework provides:

  1. Automatic Output Analysis

    • Attempts to segment output into logical blocks

    • Identifies section headers

    • Detects pass/fail indicators

    • Provides reasonable output structure even for unknown formats

  2. Flexible Integration

    • Works with arbitrary shell commands

    • No specific format requirements

    • Perfect for integration with tools like act, Docker, and custom scripts

  3. Security Features

    • Command validation to prevent harmful operations

    • Can be configured to allow specific elevated permissions when necessary

Output Format

The test runner produces structured output while preserving complete test output:

interface TestResult {
  name: string;
  passed: boolean;
  output: string[];
  rawOutput?: string;  // Complete unprocessed output
}

interface TestSummary {
  total: number;
  passed: number;
  failed: number;
  duration?: number;
}

interface ParsedResults {
  framework: string;
  tests: TestResult[];
  summary: TestSummary;
  rawOutput: string;  // Complete command output
}

Results are saved in the specified output directory:

  • test_output.log: Raw test output

  • test_errors.log: Error messages if any

  • test_results.json: Structured test results

  • summary.txt: Human-readable summary

Development

Setup

  1. Clone the repository

  2. Install dependencies:

    npm install
  3. Build the project:

    npm run build

Running Tests

npm test

The test suite includes tests for all supported frameworks and verifies both successful and failed test scenarios.

CI/CD

The project uses GitHub Actions for continuous integration:

  • Automated testing on Node.js 18.x and 20.x

  • Test results uploaded as artifacts

  • Dependabot configured for automated dependency updates

Contributing

  1. Fork the repository

  2. Create your feature branch

  3. Commit your changes

  4. Push to the branch

  5. Create a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Available Tools

1 tool
run_testsC

Run tests and capture output

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesTest command to execute (e.g., "bats tests/*.bats")
envNoEnvironment variables for test execution
frameworkYesTesting framework being used
outputDirNoDirectory to store test results
securityOptionsNoSecurity options for command execution
timeoutNoTest execution timeout in milliseconds (default: 300000)
workingDirYesWorking directory for test execution

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions 'capture output' but doesn't describe output format, error handling, side effects, or security implications. The description doesn't contradict annotations (none exist), but fails to disclose important behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just four words, front-loaded with the core action and outcome. Every word earns its place with zero redundancy or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 7 parameters, nested objects, no annotations, and no output schema, the description is insufficient. It doesn't explain return values, error conditions, security considerations, or typical usage patterns that would help an agent understand this execution tool's behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing detailed parameter documentation. The description adds no additional parameter semantics beyond the schema's comprehensive coverage, so it meets the baseline of 3 for high schema coverage without compensating value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run tests and capture output' clearly states the action (run tests) and outcome (capture output), but lacks specificity about what types of tests or how they're executed. It doesn't distinguish from siblings (none exist), but remains somewhat vague about scope and implementation details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, prerequisites, or typical scenarios. With no sibling tools mentioned, differentiation isn't needed, but there's still no context about appropriate use cases or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev1.0.0
    • First observedrun_tests

TDQS

C2.9/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as running tests and capturing output, making it distinct by default.

Naming Consistency5/5

The single tool name 'run_tests' follows a clear verb_noun pattern, which is consistent within this minimal set. There are no other tools to compare against, so no inconsistency can arise.

Tool Count2/5

A single tool is generally too few for most server purposes, as it limits functionality and flexibility. For a test runner, one tool might cover basic execution but lacks operations like listing tests, filtering, or managing test suites, making it feel thin and under-scoped.

Completeness2/5

The tool surface is severely incomplete for a test runner domain. It only provides execution without supporting operations such as listing available tests, retrieving results, configuring test runs, or handling test environments, leading to significant gaps that could cause agent failures.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    Facilitates isolated code execution within Docker containers, enabling secure multi-language script execution and integration with language models like Claude via the Model Context Protocol.
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Rocketlane's tools and services through the Model Context Protocol, enabling unified API access.
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Neon's tools and services through a unified API via the Model Context Protocol.
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/privsim/mcp-test-runner'

If you have feedback or need assistance with the MCP directory API, please join our Discord server