webdev-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@webdev-mcptake a screenshot of my current screen"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
webdev-mcp
An MCP server that provides useful web development tools.
Usage
Cursor
To install in a project, add the MCP server to your
.cursor/mcp.json:
{
"mcpServers": {
"webdev": {
"command": "npx",
"args": ["webdev-mcp"],
}
}
}To install globally, add this command to your Cursor settings:
npx webdev-mcpWindsurf
Add the MCP server to your
~/.codeium/windsurf/mcp_config.jsonfile:
{
"mcpServers": {
"webdev": {
"command": "npx",
"args": ["webdev-mcp"]
}
}
}Related MCP server: webdev-mcp
Tools
Currently, the only 2 tools are takeScreenshot and listScreens. Your agent can use the list screens tool to get the screen id of the screen it wants to screenshot.
The tool will return the screenshot as a base64 encoded string.

Tips
Make sure YOLO mode is on and MCP tools protection is off in your Cursor settings for the best experience. You might have to allow Cursor to record your screen on MacOS.
Available Tools
2 toolslistScreensB
List available screens/displays that can be captured
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the purpose but fails to disclose behavioral traits such as whether the list is real-time or cached, if it requires permissions, what format the output is in, or any rate limits. This leaves the agent with insufficient information to predict tool behavior accurately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words, front-loading the core action and purpose. Every element ('List available screens/displays that can be captured') directly contributes to understanding, making it optimally concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the return values are (e.g., list format, screen identifiers), behavioral aspects like permissions or freshness, or error handling. For a tool with zero structured metadata, this leaves significant gaps in agent understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 parameters and 100% schema description coverage, the baseline is high. The description adds value by clarifying that the screens are 'available' and for 'capture,' which provides context beyond the empty schema. However, it doesn't detail output semantics, slightly limiting its utility.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('available screens/displays'), specifying that these are for capture purposes. It distinguishes from the sibling 'takeScreenshot' by focusing on enumeration rather than action. However, it doesn't explicitly differentiate scope or limitations beyond the capture context, keeping it from a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by mentioning 'that can be captured,' suggesting this tool should be used before 'takeScreenshot' to identify targets. No explicit guidance on when not to use it or alternatives is provided, and it lacks prerequisites or error conditions, leaving gaps in decision-making context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
takeScreenshotA
Take a screenshot of a specific screen and return it as a base64 encoded string.
| Name | Required | Description | Default |
|---|---|---|---|
| screenId | No | ID of the screen to capture. Use listScreens to find available screens. Default is 1 (main screen) | |
| timeout | No | Maximum time to wait in milliseconds (default: 0, no timeout) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the output behavior (returns base64 encoded string) and implies a capture action, but lacks details on permissions needed, potential side effects (e.g., screen flicker), error conditions, or rate limits. It's adequate but has gaps for a tool that interacts with system resources.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and output. Every word earns its place, with no redundancy or unnecessary elaboration, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (capturing screenshots with two parameters), no annotations, and no output schema, the description is reasonably complete. It covers the purpose, output format, and references a sibling tool, but could benefit from more behavioral context (e.g., permissions or errors) to be fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both parameters (screenId and timeout). The description adds no additional parameter semantics beyond what's in the schema, such as format details or constraints. Baseline 3 is appropriate when the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Take a screenshot') and resource ('of a specific screen'), and distinguishes from the sibling tool 'listScreens' by mentioning it for finding available screens. It's precise about the verb, target, and output format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context by referencing the sibling tool 'listScreens' to find screen IDs, which implies when to use this tool (after identifying screens). However, it doesn't explicitly state when not to use it or name alternatives beyond the implied workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v1.0.0- Changed
listScreens1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
2 tool updates
- First observed
listScreens - First observed
takeScreenshot
TDQS
The two tools have clearly distinct purposes: listScreens identifies available screens for capture, while takeScreenshot performs the actual screenshot operation on a specified screen. There is no overlap or ambiguity between these functions, making it easy for an agent to select the correct tool based on the task.
The naming is inconsistent: listScreens uses camelCase, while takeScreenshot uses camelCase but with a different verb style (list vs. take). Although both are camelCase, the lack of a uniform verb_noun pattern (e.g., list_screens, take_screenshot) and mixed verb conventions reduce predictability and readability.
With only 2 tools, the server feels too thin for the web development domain implied by the name 'webdev-mcp'. This limited set lacks essential operations such as screen recording, window management, or other common web development tasks, making it insufficient for comprehensive coverage.
The tool surface is significantly incomplete for a web development server. While listScreens and takeScreenshot cover basic screenshot functionality, there are obvious gaps like capturing specific windows, recording screen activity, or integrating with web development workflows (e.g., browser automation, code editing). This will likely cause agent failures in broader tasks.
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Screenshot and HTML render MCP server for AI agents
MCP server for Mint — AI-powered QA that runs your app in a real browser on every PR.
Related MCP Servers
- AlicenseBqualityCmaintenanceAn official MCP server implementation that allows AI assistants to capture website screenshots through the ScreenshotOne API, enabling visual context from web pages during conversations.12636MIT
- AlicenseBqualityDmaintenanceAn MCP server providing web development tools such as screen capturing capabilities that let AI agents take and work with screenshots of the user's screen.23015MIT
- AlicenseBqualityDmaintenanceAn MCP server that enables AI assistants to capture and analyze web page screenshots using Puppeteer, supporting multi-breakpoint captures, error reporting, and page interactions.1486MIT
- AlicenseNot gradedqualityCmaintenanceA privacy-first macOS MCP server that enables AI agents to capture screenshots of pre-approved application windows for development and debugging tasks.14MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/zueai/webdev-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server