Skip to main content
Glama
dla-kirito

macos-screen-mcp

by dla-kirito

macos-screen-mcp

npm version License: MIT macOS Node

Give AI eyes on your macOS desktop — an MCP server that lets AI assistants see your screen, read browser tabs, and capture screenshots.

Features

  • Desktop awareness — frontmost app, visible apps, window positions, screen resolution

  • Screenshot capture — full screen, specific region, or frontmost window (with configurable scale)

  • Browser tab inspection — Chrome, Safari, and Arc support (active tab or all tabs)

  • File preview — open files in default app, Chrome, or Quick Look

Related MCP server: Desktop Commander MCP Server

Requirements

  • macOS 12 (Monterey) or later

  • Node.js 18+

  • Screen Recording permission (for screenshot features only)

Installation

claude mcp add --transport stdio macos-screen -- npx -y macos-screen-mcp

That's it. No global install needed.

Global install

npm install -g macos-screen-mcp
claude mcp add --transport stdio macos-screen -- macos-screen-mcp

From source

git clone https://github.com/dla-kirito/macos-screen-mcp.git
cd macos-screen-mcp
npm install && npm run build
claude mcp add --transport stdio macos-screen -- node /path/to/macos-screen-mcp/dist/index.js

Cursor / Other MCP Clients

Add to your MCP config:

{
  "mcpServers": {
    "macos-screen": {
      "command": "npx",
      "args": ["-y", "macos-screen-mcp"]
    }
  }
}

Permissions Setup

Screen Recording (required for screenshots)

The first time you use capture_screen, macOS will prompt for Screen Recording permission.

  1. Open System Settings > Privacy & Security > Screen Recording

  2. Enable the toggle for your terminal app (e.g., Ghostty, iTerm2, Terminal)

  3. Restart your terminal if prompted

Automation (for browser tools)

The first time get_desktop_state or get_browser_content reads a browser's tabs, macOS will show a dialog like "<Terminal> wants to control "Google Chrome"". This is the standard macOS Automation prompt — click OK. macOS only asks once per app pair, and you can review/revoke it later under System Settings > Privacy & Security > Automation.

Note: preview_file doesn't require any special permissions — it uses the standard open command.

Available Tools

Tool

Description

Permissions

get_desktop_state

Quick overview: frontmost app, visible apps, Chrome tabs, screen size

None

capture_screen

Screenshot (full / region / frontmost window), returns as image

Screen Recording

get_browser_content

Detailed browser tabs for Chrome, Safari, or Arc

None

preview_file

Open a file in default app, Chrome, or Quick Look

None

Tool Details

capture_screen supports a scale parameter (0–1, default 0.5) to reduce image size before sending to the LLM, saving tokens while preserving enough detail for most tasks.

get_browser_content can return just the active tab per window (default) or all tabs with include_all_tabs: true.

How It Works

The server communicates with AI clients over stdio using the Model Context Protocol. Under the hood it uses:

  • AppleScript (osascript) to query desktop state, window bounds, and browser tabs

  • screencapture (macOS built-in) to take screenshots

  • sips to downscale images before returning them as base64 PNG

All operations are read-only — the server never modifies your files, settings, or browser state.

┌─────────────┐   stdio/MCP   ┌──────────────────┐   AppleScript   ┌─────────┐
│  AI Client   │◄────────────►│  macos-screen-mcp │◄──────────────►│  macOS   │
│ (Claude etc) │              └──────────────────┘   screencapture  │ Desktop  │
└─────────────┘                                                     └─────────┘

Security

  • All tool inputs are validated via Zod schemas

  • Application names are restricted to safe characters (no shell/AppleScript injection)

  • File operations use execFile with argument arrays (no shell interpolation)

Privacy

This server runs locally and does not send data to any remote service of its own. However, by design it lets your AI assistant see parts of your desktop, and whatever the AI sees is sent to the LLM provider you've configured (Anthropic, your Cursor backend, etc.) as part of normal MCP tool responses.

What each tool exposes:

  • get_desktop_state — frontmost app, list of visible apps, Chrome window URLs and titles, window positions, screen resolution

  • get_browser_content — for the chosen browser: every window's active tab URL and title (and all tabs if include_all_tabs=true)

  • capture_screen — raw pixels of your screen / a region / the frontmost window, sent as a base64 PNG

  • preview_file — only opens the file locally; no file contents are read or transmitted by this server

Treat anything visible on screen or in a browser tab as something the AI may receive. Avoid calling these tools while password managers, private chats, banking sites, or other sensitive content are visible. Most MCP clients let you disable individual tools per session if you want a temporary lockdown.

Known Limitations

  • macOS only — relies on AppleScript and macOS-specific commands

  • Browser inspection requires the target browser to be running

  • capture_screen requires explicit Screen Recording permission

  • Arc browser only supports active tab queries (no include_all_tabs)

Contributing

git clone https://github.com/dla-kirito/macos-screen-mcp.git
cd macos-screen-mcp
npm install
npm run build    # TypeScript → dist/
npm run lint     # Type-check without emitting

License

MIT

Available Tools

4 tools
capture_screenA

Capture a screenshot of the macOS screen and return it as an image. Supports full screen, a specific region, or the frontmost window. Requires Screen Recording permission on first use.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesWhat to capture: full screen, a region, or the frontmost window
regionNoRegion to capture (only when target is 'region')
scaleNoScale factor (0-1) to reduce image size. Default 0.5

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It covers permission and target types, but omits details like output format (e.g., base64, file path), error handling, or what happens when invalid parameters are provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver the core purpose and a critical prerequisite (permission). No redundant phrases, front-loaded with main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so description should explain return format. 'Return it as an image' is vague. Lacks details on output type (e.g., base64 string, file path) and potential side effects. Adequate for a simple screen capture but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3. Description adds value by summarizing target enum and noting scale default 0.5. However, the region parameter's coordinate system (units) is not explained, and nested object constraints are implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Capture a screenshot of the macOS screen and return it as an image.' and lists three capture modes (full, region, frontmost), which distinctly separates it from sibling tools like get_browser_content or get_desktop_state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentions the required Screen Recording permission, which is a key prerequisite. It does not explicitly state when not to use, but the purpose is sufficiently distinct from siblings to infer appropriate contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_browser_contentA

Get detailed browser information: active tab title and URL for each window, optionally all tabs. Supports Chrome, Safari, and Arc. Use this when you need comprehensive browser state beyond what get_desktop_state provides.

ParametersJSON Schema
NameRequiredDescriptionDefault
browserNoBrowser to query. Default: chrome
include_all_tabsNoInclude all tabs, not just the active one. Default: false

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It accurately describes the read-only nature (getting information) and lists supported browsers. It does not mention error handling or potential permissions, but for a non-destructive retrieval, the description is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no extraneous words. It front-loads the purpose and key details, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, two-parameter, no-output-schema tool, the description covers the essential aspects: what it does, when to use it, supported browsers, and differentiation from a sibling. It does not detail the return structure or error cases, but the tool's simplicity makes this acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds context about what is returned (active tab title and URL per window, optionally all tabs) but does not significantly enhance understanding beyond the schema's parameter descriptions, which already include defaults and enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves detailed browser information (active tab title and URL for each window, optionally all tabs). It explicitly distinguishes from the sibling tool get_desktop_state, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description advises using this tool when comprehensive browser state is needed beyond get_desktop_state, providing context for when to use it. However, it lacks explicit guidance on when not to use it or what get_desktop_state specifically covers, which could be clearer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_desktop_stateA

Get the current macOS desktop context: frontmost app, visible apps, Chrome browser tabs (title + URL), window positions, and screen resolution. Use this for a quick overview of what the user is looking at.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description implies a read-only operation by listing retrieved data, but it does not explicitly state that the tool has no side effects or destructive potential. With no annotations provided, the description should disclose behavioral traits (e.g., 'read-only, no modification of state'), which is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, efficiently listing the data points and usage hint without any fluff. Every word adds value, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, no output schema, and sibling tools, the description is fairly complete: it tells what the tool retrieves and when to use it. However, it lacks mention of prerequisites (e.g., macOS accessibility permissions) or the format of the returned data, which would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is 100%. The description adds no parameter-specific meaning, but with no parameters, the baseline score of 4 is appropriate. There is no additional information needed beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (get), the resource (macOS desktop context), and lists specific elements (frontmost app, visible apps, Chrome tabs, window positions, screen resolution). It distinguishes itself from sibling tools like capture_screen, get_browser_content, and preview_file by focusing on overall desktop state rather than capturing or extracting specific content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description advises using this tool for a 'quick overview of what the user is looking at,' providing a clear context of use. However, it does not explicitly mention when not to use it or suggest alternatives among siblings, which would be helpful for an AI agent deciding between tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_fileA

Open a file for preview on macOS. Supports HTML (opens in Chrome), images, PDFs, and other files. Can use default app, Chrome, or Quick Look.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesAbsolute or relative path to the file to preview
methodNoHow to open: default app, Chrome browser, or Quick Look. Default: browser for HTML, open for others

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool opens a file for preview and offers three methods, which implies non-destructive behavior. However, it does not mention potential errors, return values, or platform dependency beyond macOS.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences, front-loading the core purpose. Every sentence adds value without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two parameters and no output schema, the description covers the main functionality (supported types, methods). It could be more complete by mentioning return behavior or error handling, but given the low complexity, it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters with descriptions. The description adds context about supported file types and method options, but these largely overlap with schema defaults. It does not introduce new meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function with a specific verb ('Open') and resource ('file for preview on macOS'), and lists supported file types. It distinguishes from sibling tools like capture_screen and get_browser_content, which are clearly different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for previewing files and mentions supported types, but does not explicitly guide when to use this tool versus alternatives or when not to use it. No reference to siblings is made.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv1.1.2
    • First observedcapture_screen
    • First observedget_browser_content
    • First observedget_desktop_state
    • First observedpreview_file

TDQS

A4/5.0
Disambiguation5/5

Each tool serves a distinct function: screen capture, detailed browser info, desktop overview, and file preview. There is no overlap in purpose.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (capture_screen, get_browser_content, get_desktop_state, preview_file).

Tool Count4/5

With 4 tools, the set is concise and well-scoped for macOS screen interaction. Slightly on the smaller side but still appropriate for the domain.

Completeness4/5

The tools cover common screen-related tasks: capturing, browser/desktop state, and file preview. Minor gaps like screen recording or window management exist but are not core to the server's stated purpose.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that gives any AI assistant eyes and hands on your desktop — screenshots, clicking, typing, OCR, window management, accessibility-tree queries, workflow recording.
    5
    Apache 2.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/dla-kirito/macos-screen-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server