macos-screen-mcp
Allows reading active browser tabs from Arc (limited to active tab only).
Allows reading active and all browser tabs from Google Chrome, including URLs and titles.
Provides desktop state awareness (frontmost app, visible apps, window positions, screen resolution) and screenshot capture (full screen, region, or frontmost window).
Allows reading active and all browser tabs from Safari, including URLs and titles.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@macos-screen-mcpwhat's on my screen?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
macos-screen-mcp
Give AI eyes on your macOS desktop — an MCP server that lets AI assistants see your screen, read browser tabs, and capture screenshots.
Features
Desktop awareness — frontmost app, visible apps, window positions, screen resolution
Screenshot capture — full screen, specific region, or frontmost window (with configurable scale)
Browser tab inspection — Chrome, Safari, and Arc support (active tab or all tabs)
File preview — open files in default app, Chrome, or Quick Look
Related MCP server: Desktop Commander MCP Server
Requirements
macOS 12 (Monterey) or later
Node.js 18+
Screen Recording permission (for screenshot features only)
Installation
Quick start (recommended)
claude mcp add --transport stdio macos-screen -- npx -y macos-screen-mcpThat's it. No global install needed.
Global install
npm install -g macos-screen-mcp
claude mcp add --transport stdio macos-screen -- macos-screen-mcpFrom source
git clone https://github.com/dla-kirito/macos-screen-mcp.git
cd macos-screen-mcp
npm install && npm run build
claude mcp add --transport stdio macos-screen -- node /path/to/macos-screen-mcp/dist/index.jsCursor / Other MCP Clients
Add to your MCP config:
{
"mcpServers": {
"macos-screen": {
"command": "npx",
"args": ["-y", "macos-screen-mcp"]
}
}
}Permissions Setup
Screen Recording (required for screenshots)
The first time you use capture_screen, macOS will prompt for Screen Recording permission.
Open System Settings > Privacy & Security > Screen Recording
Enable the toggle for your terminal app (e.g., Ghostty, iTerm2, Terminal)
Restart your terminal if prompted
Automation (for browser tools)
The first time get_desktop_state or get_browser_content reads a browser's tabs, macOS will show a dialog like "<Terminal> wants to control "Google Chrome"". This is the standard macOS Automation prompt — click OK. macOS only asks once per app pair, and you can review/revoke it later under System Settings > Privacy & Security > Automation.
Note:
preview_filedoesn't require any special permissions — it uses the standardopencommand.
Available Tools
Tool | Description | Permissions |
| Quick overview: frontmost app, visible apps, Chrome tabs, screen size | None |
| Screenshot (full / region / frontmost window), returns as image | Screen Recording |
| Detailed browser tabs for Chrome, Safari, or Arc | None |
| Open a file in default app, Chrome, or Quick Look | None |
Tool Details
capture_screen supports a scale parameter (0–1, default 0.5) to reduce image size before sending to the LLM, saving tokens while preserving enough detail for most tasks.
get_browser_content can return just the active tab per window (default) or all tabs with include_all_tabs: true.
How It Works
The server communicates with AI clients over stdio using the Model Context Protocol. Under the hood it uses:
AppleScript (
osascript) to query desktop state, window bounds, and browser tabsscreencapture(macOS built-in) to take screenshotssipsto downscale images before returning them as base64 PNG
All operations are read-only — the server never modifies your files, settings, or browser state.
┌─────────────┐ stdio/MCP ┌──────────────────┐ AppleScript ┌─────────┐
│ AI Client │◄────────────►│ macos-screen-mcp │◄──────────────►│ macOS │
│ (Claude etc) │ └──────────────────┘ screencapture │ Desktop │
└─────────────┘ └─────────┘Security
All tool inputs are validated via Zod schemas
Application names are restricted to safe characters (no shell/AppleScript injection)
File operations use
execFilewith argument arrays (no shell interpolation)
Privacy
This server runs locally and does not send data to any remote service of its own. However, by design it lets your AI assistant see parts of your desktop, and whatever the AI sees is sent to the LLM provider you've configured (Anthropic, your Cursor backend, etc.) as part of normal MCP tool responses.
What each tool exposes:
get_desktop_state— frontmost app, list of visible apps, Chrome window URLs and titles, window positions, screen resolutionget_browser_content— for the chosen browser: every window's active tab URL and title (and all tabs ifinclude_all_tabs=true)capture_screen— raw pixels of your screen / a region / the frontmost window, sent as a base64 PNGpreview_file— only opens the file locally; no file contents are read or transmitted by this server
Treat anything visible on screen or in a browser tab as something the AI may receive. Avoid calling these tools while password managers, private chats, banking sites, or other sensitive content are visible. Most MCP clients let you disable individual tools per session if you want a temporary lockdown.
Known Limitations
macOS only — relies on AppleScript and macOS-specific commands
Browser inspection requires the target browser to be running
capture_screenrequires explicit Screen Recording permissionArc browser only supports active tab queries (no
include_all_tabs)
Contributing
git clone https://github.com/dla-kirito/macos-screen-mcp.git
cd macos-screen-mcp
npm install
npm run build # TypeScript → dist/
npm run lint # Type-check without emittingLicense
Available Tools
4 toolscapture_screenA
Capture a screenshot of the macOS screen and return it as an image. Supports full screen, a specific region, or the frontmost window. Requires Screen Recording permission on first use.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | What to capture: full screen, a region, or the frontmost window | |
| region | No | Region to capture (only when target is 'region') | |
| scale | No | Scale factor (0-1) to reduce image size. Default 0.5 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It covers permission and target types, but omits details like output format (e.g., base64, file path), error handling, or what happens when invalid parameters are provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver the core purpose and a critical prerequisite (permission). No redundant phrases, front-loaded with main action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so description should explain return format. 'Return it as an image' is vague. Lacks details on output type (e.g., base64 string, file path) and potential side effects. Adequate for a simple screen capture but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so baseline is 3. Description adds value by summarizing target enum and noting scale default 0.5. However, the region parameter's coordinate system (units) is not explained, and nested object constraints are implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Capture a screenshot of the macOS screen and return it as an image.' and lists three capture modes (full, region, frontmost), which distinctly separates it from sibling tools like get_browser_content or get_desktop_state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions the required Screen Recording permission, which is a key prerequisite. It does not explicitly state when not to use, but the purpose is sufficiently distinct from siblings to infer appropriate contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_browser_contentA
Get detailed browser information: active tab title and URL for each window, optionally all tabs. Supports Chrome, Safari, and Arc. Use this when you need comprehensive browser state beyond what get_desktop_state provides.
| Name | Required | Description | Default |
|---|---|---|---|
| browser | No | Browser to query. Default: chrome | |
| include_all_tabs | No | Include all tabs, not just the active one. Default: false |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It accurately describes the read-only nature (getting information) and lists supported browsers. It does not mention error handling or potential permissions, but for a non-destructive retrieval, the description is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no extraneous words. It front-loads the purpose and key details, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, two-parameter, no-output-schema tool, the description covers the essential aspects: what it does, when to use it, supported browsers, and differentiation from a sibling. It does not detail the return structure or error cases, but the tool's simplicity makes this acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds context about what is returned (active tab title and URL per window, optionally all tabs) but does not significantly enhance understanding beyond the schema's parameter descriptions, which already include defaults and enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves detailed browser information (active tab title and URL for each window, optionally all tabs). It explicitly distinguishes from the sibling tool get_desktop_state, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using this tool when comprehensive browser state is needed beyond get_desktop_state, providing context for when to use it. However, it lacks explicit guidance on when not to use it or what get_desktop_state specifically covers, which could be clearer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_desktop_stateA
Get the current macOS desktop context: frontmost app, visible apps, Chrome browser tabs (title + URL), window positions, and screen resolution. Use this for a quick overview of what the user is looking at.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description implies a read-only operation by listing retrieved data, but it does not explicitly state that the tool has no side effects or destructive potential. With no annotations provided, the description should disclose behavioral traits (e.g., 'read-only, no modification of state'), which is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, efficiently listing the data points and usage hint without any fluff. Every word adds value, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters, no output schema, and sibling tools, the description is fairly complete: it tells what the tool retrieves and when to use it. However, it lacks mention of prerequisites (e.g., macOS accessibility permissions) or the format of the returned data, which would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is 100%. The description adds no parameter-specific meaning, but with no parameters, the baseline score of 4 is appropriate. There is no additional information needed beyond what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (get), the resource (macOS desktop context), and lists specific elements (frontmost app, visible apps, Chrome tabs, window positions, screen resolution). It distinguishes itself from sibling tools like capture_screen, get_browser_content, and preview_file by focusing on overall desktop state rather than capturing or extracting specific content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using this tool for a 'quick overview of what the user is looking at,' providing a clear context of use. However, it does not explicitly mention when not to use it or suggest alternatives among siblings, which would be helpful for an AI agent deciding between tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_fileA
Open a file for preview on macOS. Supports HTML (opens in Chrome), images, PDFs, and other files. Can use default app, Chrome, or Quick Look.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Absolute or relative path to the file to preview | |
| method | No | How to open: default app, Chrome browser, or Quick Look. Default: browser for HTML, open for others |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool opens a file for preview and offers three methods, which implies non-destructive behavior. However, it does not mention potential errors, return values, or platform dependency beyond macOS.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences, front-loading the core purpose. Every sentence adds value without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters and no output schema, the description covers the main functionality (supported types, methods). It could be more complete by mentioning return behavior or error handling, but given the low complexity, it is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of parameters with descriptions. The description adds context about supported file types and method options, but these largely overlap with schema defaults. It does not introduce new meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Open') and resource ('file for preview on macOS'), and lists supported file types. It distinguishes from sibling tools like capture_screen and get_browser_content, which are clearly different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for previewing files and mentions supported types, but does not explicitly guide when to use this tool versus alternatives or when not to use it. No reference to siblings is made.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v1.1.2- First observed
capture_screen - First observed
get_browser_content - First observed
get_desktop_state - First observed
preview_file
TDQS
Each tool serves a distinct function: screen capture, detailed browser info, desktop overview, and file preview. There is no overlap in purpose.
All tools follow a consistent snake_case verb_noun pattern (capture_screen, get_browser_content, get_desktop_state, preview_file).
With 4 tools, the set is concise and well-scoped for macOS screen interaction. Slightly on the smaller side but still appropriate for the domain.
The tools cover common screen-related tasks: capturing, browser/desktop state, and file preview. Minor gaps like screen recording or window management exist but are not core to the server's stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Driflyte MCP server which lets AI assistants query topic-specific knowledge from web and GitHub.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
An MCP server that gives your AI access to the source code and docs of all public github repos
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that allows AI tools like Claude Desktop, Claude Code, and Cursor to visually interact with macOS applications by capturing screenshots and controlling the mouse and keyboard.14-
- AlicenseNot gradedqualityDmaintenanceA comprehensive MCP server that gives AI assistants full control over your desktop — monitor system resources, manage windows, capture screenshots, control the clipboard, launch applications, and more.MIT
- AlicenseNot gradedqualityBmaintenanceAn MCP server that gives any AI assistant eyes and hands on your desktop — screenshots, clicking, typing, OCR, window management, accessibility-tree queries, workflow recording.5Apache 2.0
- AlicenseNot gradedqualityDmaintenanceAn open-source MCP server for macOS that lets AI control desktop apps in the background without moving the cursor or stealing focus.25Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/dla-kirito/macos-screen-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server