Skip to main content
Glama
taskhub-sh

Terminal Control MCP

by taskhub-sh

Terminal Control MCP

A Model Context Protocol (MCP) server that enables AI agents to interact with terminal-based TUI applications through a virtual X11 display approach.

Overview

This project provides a comprehensive solution for controlling terminal applications programmatically using:

  • Xvfb for headless virtual X11 display

  • xterm for terminal emulation

  • xdotool for input simulation and window management

  • ImageMagick for PNG screenshot capture

The system captures actual visual terminal output as PNG screenshots, making it ideal for AI agents that need to see and interact with terminal applications.

Related MCP server: Interactive Terminal MCP Server

Features

  • Virtual Display Management: Headless X11 display using Xvfb

  • Input Simulation: Send keyboard input and text to terminal applications

  • Screenshot Capture: Take PNG screenshots of terminal output

  • Window Management: Reliable window detection and focus handling

  • Resource Cleanup: Proper process management with timeout handling

System Requirements

The following system packages must be installed:

# Ubuntu/Debian
sudo apt-get install xvfb xterm xdotool imagemagick

# CentOS/RHEL/Fedora
sudo yum install xorg-x11-server-Xvfb xterm xdotool ImageMagick

Installation

This project uses the uv package manager:

# Clone the repository
git clone <repository-url>
cd terminal-control-mcp

# Install dependencies
uv sync

# Activate virtual environment
source .venv/bin/activate

Quick Start

Running the Example

Try the included htop example to see the system in action:

python examples/example_htop.py

This will:

  1. Launch htop in a virtual xterm session

  2. Press F3 to open the search dialog

  3. Type "python" as a search term

  4. Capture PNG screenshots at each step

  5. Clean up all processes

Basic Usage

from examples.example_htop import XTermSession

# Create a session
session = XTermSession(width=1920, height=1080)

try:
    # Start virtual display and terminal
    session.start_virtual_display()
    session.start_xterm("your-command-here")
    
    # Take a screenshot
    session.take_screenshot("output.png")
    
    # Send input
    session.send_key("F1")
    session.send_text("hello world")
    
finally:
    session.cleanup()

Architecture

Core Components

  1. XTermSession Class: The main interface for terminal control

    • Manages Xvfb virtual display lifecycle

    • Spawns and controls xterm processes

    • Handles input simulation via xdotool

    • Captures screenshots using ImageMagick

  2. Virtual Display Approach: Unlike direct TTY manipulation, this system:

    • Creates a real X11 environment with Xvfb

    • Launches actual xterm instances

    • Captures genuine visual output as PNG files

    • Provides reliable input simulation

Key Methods

  • start_virtual_display(): Initialize Xvfb virtual display

  • start_xterm(command): Launch xterm with specified command

  • send_key(key): Send special keys (F1, Escape, etc.)

  • send_text(text): Send alphanumeric text input

  • take_screenshot(filename): Capture PNG screenshot

  • cleanup(): Properly terminate all processes

Development

Project Structure

terminal-control-mcp/
├── src/terminal_control_mcp/    # Main MCP server implementation (planned)
├── examples/
│   ├── example_htop.py         # Reference implementation
│   └── README.md
├── tests/                      # Test suite
├── pyproject.toml             # Project configuration
└── CLAUDE.md                  # Development guidelines

Development Commands

# Run the main application
python main.py

# Run the htop example
python examples/example_htop.py

# Activate virtual environment
source .venv/bin/activate

MCP Server Implementation (Planned)

The full MCP server will provide these tools:

  • terminal_launch: Start a new terminal session

  • terminal_input: Send keyboard/text input

  • terminal_capture: Take PNG screenshot

  • terminal_close: Clean up terminal session

Technical Details

Window Detection

The system uses multiple fallback strategies for reliable window ID detection:

search_methods = [
    ['xdotool', 'search', '--class', 'XTerm'],
    ['xdotool', 'search', '--name', 'xterm'],
    ['xdotool', 'search', '--class', 'xterm'],
    ['xdotool', 'getactivewindow']
]

Screenshot Capture

Uses ImageMagick's import command for reliable PNG capture:

subprocess.run(['import', '-window', 'root', filename], env=env)

Resource Management

Implements proper cleanup with timeout handling:

def cleanup(self):
    if self.xterm_proc:
        self.xterm_proc.terminate()
        try:
            self.xterm_proc.wait(timeout=5)
        except subprocess.TimeoutExpired:
            self.xterm_proc.kill()

Contributing

[TBD]

Available Tools

4 tools
terminal_captureB

Capture terminal screen as PNG screenshot

Args: session_id: ID of the terminal session

Returns: Dictionary with base64-encoded PNG image and metadata

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool captures a screenshot and returns a dictionary with base64-encoded PNG and metadata, but it does not cover critical aspects like required permissions, rate limits, error handling, or whether the capture is instantaneous or delayed. For a tool with no annotations, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core action and following with structured sections for Args and Returns. Every sentence earns its place by conveying key information without redundancy, though the formatting could be slightly more polished for optimal clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (single parameter, no annotations, but with an output schema), the description is partially complete. It covers the basic purpose, parameter semantics, and return structure, but lacks usage guidelines and behavioral details. The presence of an output schema reduces the need to explain return values, but overall, it falls short of being fully comprehensive for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful context for the single parameter 'session_id' by specifying it as 'ID of the terminal session,' which clarifies its purpose beyond the schema's minimal title 'Session Id.' With 0% schema description coverage and only one parameter, the description effectively compensates by providing essential semantic information, though it could elaborate on format or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Capture terminal screen as PNG screenshot') and resource ('terminal session'), distinguishing it from sibling tools like terminal_close, terminal_input, and terminal_launch. It precisely communicates what the tool does without being vague or tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. While it implies usage for capturing screenshots of terminal sessions, it does not mention prerequisites, exclusions, or specific scenarios where this tool is preferred over others. No explicit when/when-not or alternative tool references are included.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminal_closeA

Close a terminal session and cleanup resources

Args: session_id: ID of the terminal session

Returns: Dictionary with cleanup status

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the action ('close') and a behavioral trait ('cleanup resources'), which implies resource management, but does not detail what cleanup entails (e.g., freeing memory, terminating processes), potential side effects, or error conditions. It adds some value but lacks depth for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by structured sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy, making it efficient and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (a mutation with cleanup), no annotations, and an output schema (which handles return values), the description is reasonably complete. It covers purpose, parameters, and returns, but could improve by adding more behavioral context (e.g., what happens on failure) to fully compensate for the lack of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It adds meaning by explaining that 'session_id' is the 'ID of the terminal session', which clarifies its role beyond the schema's basic type. However, it does not provide format details, examples, or constraints, leaving gaps in parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Close a terminal session and cleanup resources') with the target resource ('terminal session'), distinguishing it from sibling tools like terminal_launch (create), terminal_input (send input), and terminal_capture (capture output). The verb 'close' and resource 'terminal session' are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying 'close a terminal session', suggesting it should be used when a session is no longer needed, but it does not explicitly state when not to use it or name alternatives. The context is clear but lacks explicit exclusions or comparisons to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminal_inputA

Send input to a terminal session

Args: session_id: ID of the terminal session input_text: Text to type (for alphanumeric input) key: Special key to send (Return, Tab, Escape, etc.)

Returns: Dictionary with status and input confirmation

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes
input_textNo
keyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions sending input and a return dictionary, but lacks critical behavioral details such as whether this is a read-only or mutating operation, error handling (e.g., invalid session_id), rate limits, or how input_text and key interact (e.g., if both are provided). The description doesn't disclose enough about the tool's behavior beyond basic functionality.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear purpose statement followed by Args and Returns sections. It's appropriately sized and front-loaded, with each sentence adding value. Minor improvements could include integrating the sections more fluidly, but overall it's efficient with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 parameters, no annotations, but has an output schema), the description is fairly complete. It covers the purpose, parameters, and return value. The output schema exists, so the description doesn't need to detail return values, but it could improve by addressing behavioral aspects like error cases or interaction between input_text and key.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by explaining all three parameters: session_id (ID of the terminal session), input_text (text to type for alphanumeric input), and key (special key like Return, Tab, Escape). It adds meaningful context beyond the schema's titles, clarifying the purpose and usage of each parameter, though it could detail allowed values for 'key' more explicitly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Send input to a terminal session') and resource ('terminal session'), distinguishing it from siblings like terminal_capture (capture output), terminal_close (end session), and terminal_launch (start session). The verb 'send input' precisely defines the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by mentioning 'terminal session' and parameters like 'input_text' and 'key', but doesn't explicitly state when to use this tool versus alternatives. For example, it doesn't clarify if this should be used for all terminal input or only specific scenarios, nor does it mention prerequisites like needing an active session from terminal_launch.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminal_launchA

Launch a new terminal session with virtual X11 display

Args: command: Command to run in terminal (default: bash) width: Terminal width in characters (default: 80) height: Terminal height in characters (default: 24)

Returns: Dictionary with session_id and status

ParametersJSON Schema
NameRequiredDescriptionDefault
commandNobash
widthNo
heightNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses key behavioral traits: it creates a new session, uses virtual X11 display, and returns session_id and status. However, it doesn't mention important aspects like whether sessions persist, authentication requirements, resource limits, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Perfectly structured with a clear purpose statement followed by organized sections for Args and Returns. Every sentence earns its place, with no redundant information. The description is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (returns dictionary with session_id and status), the description doesn't need to explain return values. It covers the core functionality well but could benefit from more behavioral context about session lifecycle, especially since no annotations are provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate. It provides meaningful context for all 3 parameters: command specifies what to run, width and height define terminal dimensions. The default values are clearly stated, though it doesn't explain parameter constraints or valid ranges.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Launch a new terminal session') and distinguishes it from siblings like terminal_capture, terminal_close, and terminal_input by specifying it creates a new session with virtual X11 display rather than interacting with existing sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool (to start a new terminal session), but doesn't explicitly state when not to use it or mention alternatives like using terminal_input for existing sessions. The sibling tool names provide implicit differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updates
    • First observedterminal_capture
    • First observedterminal_close
    • First observedterminal_input
    • First observedterminal_launch

TDQS

A4/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: launch creates a session, input sends keystrokes, capture takes screenshots, and close terminates sessions. The descriptions reinforce these distinct roles, making it easy for an agent to select the right tool for any terminal control task.

Naming Consistency5/5

All tool names follow a perfect verb_noun pattern with 'terminal_' prefix (terminal_launch, terminal_input, terminal_capture, terminal_close). This consistent naming convention makes the tool set predictable and easy to understand at a glance.

Tool Count5/5

Four tools is ideal for this server's purpose of terminal control. Each tool earns its place by covering essential lifecycle operations: create (launch), interact (input), monitor (capture), and destroy (close). No tool feels redundant or missing for basic terminal management.

Completeness5/5

The tool set provides complete CRUD/lifecycle coverage for terminal sessions: launch creates, input interacts, capture monitors, and close deletes. There are no obvious gaps—agents can fully manage terminal sessions from creation to termination with all necessary operations available.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides AI agents with fully interactive terminal sessions, including TUI support, keyboard control, and screen capture across Windows, Linux, and Mac.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to interact with interactive CLI processes via a real PTY, allowing them to send keystrokes, read screen output, and handle interactive prompts.
    6
    1
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    Enables AI assistants to test Terminal User Interface (TUI) applications by launching, interacting with, and verifying programmatic output and behavior.
    15
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/taskhub-sh/terminal-driver-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server