Terminal Control MCP
Provides capabilities to control terminal-based applications on CentOS systems through a virtual X11 display, supporting input simulation and screenshot capture.
Enables interaction with terminal-based applications on Debian systems through virtual display management, input simulation, and screenshot capture capabilities.
Supports controlling terminal applications on Fedora systems using virtual X11 display, allowing for keyboard input simulation and visual output capture.
Enables AI agents to interact with the htop process monitoring tool, including launching htop sessions, navigating through the interface, using search functionality, and capturing visual output as PNG screenshots.
Allows AI agents to interact with terminal-based applications on Ubuntu systems through virtual display management, input simulation, and screenshot capture.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Terminal Control MCPlaunch htop and take a screenshot of the process list"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Terminal Control MCP
A Model Context Protocol (MCP) server that enables AI agents to interact with terminal-based TUI applications through a virtual X11 display approach.
Overview
This project provides a comprehensive solution for controlling terminal applications programmatically using:
Xvfb for headless virtual X11 display
xterm for terminal emulation
xdotool for input simulation and window management
ImageMagick for PNG screenshot capture
The system captures actual visual terminal output as PNG screenshots, making it ideal for AI agents that need to see and interact with terminal applications.
Related MCP server: Interactive Terminal MCP Server
Features
Virtual Display Management: Headless X11 display using Xvfb
Input Simulation: Send keyboard input and text to terminal applications
Screenshot Capture: Take PNG screenshots of terminal output
Window Management: Reliable window detection and focus handling
Resource Cleanup: Proper process management with timeout handling
System Requirements
The following system packages must be installed:
# Ubuntu/Debian
sudo apt-get install xvfb xterm xdotool imagemagick
# CentOS/RHEL/Fedora
sudo yum install xorg-x11-server-Xvfb xterm xdotool ImageMagickInstallation
This project uses the uv package manager:
# Clone the repository
git clone <repository-url>
cd terminal-control-mcp
# Install dependencies
uv sync
# Activate virtual environment
source .venv/bin/activateQuick Start
Running the Example
Try the included htop example to see the system in action:
python examples/example_htop.pyThis will:
Launch htop in a virtual xterm session
Press F3 to open the search dialog
Type "python" as a search term
Capture PNG screenshots at each step
Clean up all processes
Basic Usage
from examples.example_htop import XTermSession
# Create a session
session = XTermSession(width=1920, height=1080)
try:
# Start virtual display and terminal
session.start_virtual_display()
session.start_xterm("your-command-here")
# Take a screenshot
session.take_screenshot("output.png")
# Send input
session.send_key("F1")
session.send_text("hello world")
finally:
session.cleanup()Architecture
Core Components
XTermSession Class: The main interface for terminal control
Manages Xvfb virtual display lifecycle
Spawns and controls xterm processes
Handles input simulation via xdotool
Captures screenshots using ImageMagick
Virtual Display Approach: Unlike direct TTY manipulation, this system:
Creates a real X11 environment with Xvfb
Launches actual xterm instances
Captures genuine visual output as PNG files
Provides reliable input simulation
Key Methods
start_virtual_display(): Initialize Xvfb virtual displaystart_xterm(command): Launch xterm with specified commandsend_key(key): Send special keys (F1, Escape, etc.)send_text(text): Send alphanumeric text inputtake_screenshot(filename): Capture PNG screenshotcleanup(): Properly terminate all processes
Development
Project Structure
terminal-control-mcp/
├── src/terminal_control_mcp/ # Main MCP server implementation (planned)
├── examples/
│ ├── example_htop.py # Reference implementation
│ └── README.md
├── tests/ # Test suite
├── pyproject.toml # Project configuration
└── CLAUDE.md # Development guidelinesDevelopment Commands
# Run the main application
python main.py
# Run the htop example
python examples/example_htop.py
# Activate virtual environment
source .venv/bin/activateMCP Server Implementation (Planned)
The full MCP server will provide these tools:
terminal_launch: Start a new terminal sessionterminal_input: Send keyboard/text inputterminal_capture: Take PNG screenshotterminal_close: Clean up terminal session
Technical Details
Window Detection
The system uses multiple fallback strategies for reliable window ID detection:
search_methods = [
['xdotool', 'search', '--class', 'XTerm'],
['xdotool', 'search', '--name', 'xterm'],
['xdotool', 'search', '--class', 'xterm'],
['xdotool', 'getactivewindow']
]Screenshot Capture
Uses ImageMagick's import command for reliable PNG capture:
subprocess.run(['import', '-window', 'root', filename], env=env)Resource Management
Implements proper cleanup with timeout handling:
def cleanup(self):
if self.xterm_proc:
self.xterm_proc.terminate()
try:
self.xterm_proc.wait(timeout=5)
except subprocess.TimeoutExpired:
self.xterm_proc.kill()Contributing
[TBD]
Available Tools
4 toolsterminal_captureB
Capture terminal screen as PNG screenshot
Args: session_id: ID of the terminal session
Returns: Dictionary with base64-encoded PNG image and metadata
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool captures a screenshot and returns a dictionary with base64-encoded PNG and metadata, but it does not cover critical aspects like required permissions, rate limits, error handling, or whether the capture is instantaneous or delayed. For a tool with no annotations, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core action and following with structured sections for Args and Returns. Every sentence earns its place by conveying key information without redundancy, though the formatting could be slightly more polished for optimal clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (single parameter, no annotations, but with an output schema), the description is partially complete. It covers the basic purpose, parameter semantics, and return structure, but lacks usage guidelines and behavioral details. The presence of an output schema reduces the need to explain return values, but overall, it falls short of being fully comprehensive for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful context for the single parameter 'session_id' by specifying it as 'ID of the terminal session,' which clarifies its purpose beyond the schema's minimal title 'Session Id.' With 0% schema description coverage and only one parameter, the description effectively compensates by providing essential semantic information, though it could elaborate on format or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Capture terminal screen as PNG screenshot') and resource ('terminal session'), distinguishing it from sibling tools like terminal_close, terminal_input, and terminal_launch. It precisely communicates what the tool does without being vague or tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. While it implies usage for capturing screenshots of terminal sessions, it does not mention prerequisites, exclusions, or specific scenarios where this tool is preferred over others. No explicit when/when-not or alternative tool references are included.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terminal_closeA
Close a terminal session and cleanup resources
Args: session_id: ID of the terminal session
Returns: Dictionary with cleanup status
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the action ('close') and a behavioral trait ('cleanup resources'), which implies resource management, but does not detail what cleanup entails (e.g., freeing memory, terminating processes), potential side effects, or error conditions. It adds some value but lacks depth for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by structured sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy, making it efficient and well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a mutation with cleanup), no annotations, and an output schema (which handles return values), the description is reasonably complete. It covers purpose, parameters, and returns, but could improve by adding more behavioral context (e.g., what happens on failure) to fully compensate for the lack of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate. It adds meaning by explaining that 'session_id' is the 'ID of the terminal session', which clarifies its role beyond the schema's basic type. However, it does not provide format details, examples, or constraints, leaving gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Close a terminal session and cleanup resources') with the target resource ('terminal session'), distinguishing it from sibling tools like terminal_launch (create), terminal_input (send input), and terminal_capture (capture output). The verb 'close' and resource 'terminal session' are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying 'close a terminal session', suggesting it should be used when a session is no longer needed, but it does not explicitly state when not to use it or name alternatives. The context is clear but lacks explicit exclusions or comparisons to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terminal_inputA
Send input to a terminal session
Args: session_id: ID of the terminal session input_text: Text to type (for alphanumeric input) key: Special key to send (Return, Tab, Escape, etc.)
Returns: Dictionary with status and input confirmation
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | ||
| input_text | No | ||
| key | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions sending input and a return dictionary, but lacks critical behavioral details such as whether this is a read-only or mutating operation, error handling (e.g., invalid session_id), rate limits, or how input_text and key interact (e.g., if both are provided). The description doesn't disclose enough about the tool's behavior beyond basic functionality.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear purpose statement followed by Args and Returns sections. It's appropriately sized and front-loaded, with each sentence adding value. Minor improvements could include integrating the sections more fluidly, but overall it's efficient with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no annotations, but has an output schema), the description is fairly complete. It covers the purpose, parameters, and return value. The output schema exists, so the description doesn't need to detail return values, but it could improve by addressing behavioral aspects like error cases or interaction between input_text and key.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining all three parameters: session_id (ID of the terminal session), input_text (text to type for alphanumeric input), and key (special key like Return, Tab, Escape). It adds meaningful context beyond the schema's titles, clarifying the purpose and usage of each parameter, though it could detail allowed values for 'key' more explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Send input to a terminal session') and resource ('terminal session'), distinguishing it from siblings like terminal_capture (capture output), terminal_close (end session), and terminal_launch (start session). The verb 'send input' precisely defines the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning 'terminal session' and parameters like 'input_text' and 'key', but doesn't explicitly state when to use this tool versus alternatives. For example, it doesn't clarify if this should be used for all terminal input or only specific scenarios, nor does it mention prerequisites like needing an active session from terminal_launch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terminal_launchA
Launch a new terminal session with virtual X11 display
Args: command: Command to run in terminal (default: bash) width: Terminal width in characters (default: 80) height: Terminal height in characters (default: 24)
Returns: Dictionary with session_id and status
| Name | Required | Description | Default |
|---|---|---|---|
| command | No | bash | |
| width | No | ||
| height | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavioral traits: it creates a new session, uses virtual X11 display, and returns session_id and status. However, it doesn't mention important aspects like whether sessions persist, authentication requirements, resource limits, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Perfectly structured with a clear purpose statement followed by organized sections for Args and Returns. Every sentence earns its place, with no redundant information. The description is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (returns dictionary with session_id and status), the description doesn't need to explain return values. It covers the core functionality well but could benefit from more behavioral context about session lifecycle, especially since no annotations are provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It provides meaningful context for all 3 parameters: command specifies what to run, width and height define terminal dimensions. The default values are clearly stated, though it doesn't explain parameter constraints or valid ranges.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Launch a new terminal session') and distinguishes it from siblings like terminal_capture, terminal_close, and terminal_input by specifying it creates a new session with virtual X11 display rather than interacting with existing sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (to start a new terminal session), but doesn't explicitly state when not to use it or mention alternatives like using terminal_input for existing sessions. The sibling tool names provide implicit differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
- First observed
terminal_capture - First observed
terminal_close - First observed
terminal_input - First observed
terminal_launch
TDQS
Each tool has a clearly distinct purpose with no overlap: launch creates a session, input sends keystrokes, capture takes screenshots, and close terminates sessions. The descriptions reinforce these distinct roles, making it easy for an agent to select the right tool for any terminal control task.
All tool names follow a perfect verb_noun pattern with 'terminal_' prefix (terminal_launch, terminal_input, terminal_capture, terminal_close). This consistent naming convention makes the tool set predictable and easy to understand at a glance.
Four tools is ideal for this server's purpose of terminal control. Each tool earns its place by covering essential lifecycle operations: create (launch), interact (input), monitor (capture), and destroy (close). No tool feels redundant or missing for basic terminal management.
The tool set provides complete CRUD/lifecycle coverage for terminal sessions: launch creates, input interacts, capture monitors, and close deletes. There are no obvious gaps—agents can fully manage terminal sessions from creation to termination with all necessary operations available.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate images, GIFs, and PDFs from HTML, URLs, or templates — from your AI agent.
Run AI customer support from your terminal: conversations, knowledge base, and chat widget.
Docs for agent-manager, the terminal UI that runs AI coding agents as live tmux sessions.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
61
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server that lets AI agents see and interact with terminal/CLI applications through virtual terminals and PNG screenshots.153MIT
- AlicenseNot gradedqualityCmaintenanceProvides AI agents with fully interactive terminal sessions, including TUI support, keyboard control, and screen capture across Windows, Linux, and Mac.MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to interact with interactive CLI processes via a real PTY, allowing them to send keystrokes, read screen output, and handle interactive prompts.61MIT
- AlicenseNot gradedqualityFmaintenanceEnables AI assistants to test Terminal User Interface (TUI) applications by launching, interacting with, and verifying programmatic output and behavior.15MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/taskhub-sh/terminal-driver-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server