Enables MCP clients to discover desktop applications with structured metadata and receive native accessibility trees and screenshot content through Open Computer Use, preserving upstream tool results and image blocks.
An open-source MCP server that lets any coding agent operate your computer like a person does—reading the UI through accessibility trees, clicking and typing in the background, showing an agent pointer, zooming into regions, and driving your signed-in Chrome on macOS, Windows, and Linux.
Exposes Anthropic's computer-use action surface (screenshot, click, move, keyboard, clipboard, batch) against a persistent desktop display via MCP stdio protocol. Enables AI agents to control a virtual desktop environment through natural language instructions.
Enables safe automation of Chrome browser through a local MCP server and Chrome extension, allowing LLMs to control browser tabs, pages, and computer-use actions with permission controls.
Enables localized macOS control by letting users grant temporary, window-scoped access for screen capture, pointer, keyboard, scrolling, clipboard, and Accessibility actions through a native host.
A framework-agnostic computer-use MCP server that exposes core desktop operations (screen capture, mouse, keyboard, and file access) as standard MCP tools, enabling any MCP-compatible agent to drive a computer.
Enables Windows desktop automation via MCP, allowing AI agents to control mouse, keyboard, and screen capture with the same interface as Anthropic's computer-use tool.
Standalone MCP server that gives AI agents full GUI control over macOS — screenshots, mouse, keyboard, apps, clipboard, and multi-display — with zero private dependencies.
An ultra-fast native MCP server for macOS desktop automation, enabling visual OCR, text-based clicking, window management, and keyboard/mouse control without hijacking the physical cursor.
Enables AI agents to see, locate UI elements, and operate any Windows desktop app through natural language, using accessibility-tree matching with optional vision-model fallback, plus an autonomous visual loop with introspection and meta-learning.
Enables the Pi coding agent to drive Chrome through DevTools, providing browser automation tools, persistent or isolated profiles, screenshots, page analysis, and bundled agent policies.
Provides a JSON-RPC computer use runtime for macOS, exposing 7 MCP tools (observe/act/inspect/session/cancel/trace) as image content blocks so external agents like Claude Code, Pi, OpenCode, or Codex CLI can capture screenshots and drive the desktop with clicks, keys, and typing while enforcing session locking, stale-frame protection, and trace redaction server-side.
Enables MCP-speaking clients to control a real desktop via the computer_use tool, including clicking, typing, scrolling, dragging, key combos, app focus, and screen/accessibility capture. It also provides a verdict system that verifies whether input actions had their intended effect, with configurable approval modes.
Provides safe Linux desktop automation for Codex through focused MCP tools for inspecting, clicking, typing, dragging, and verifying native applications on Wayland and X11.