MCP Webcam Server
The MCP Webcam Server allows you to interact with your webcam and screen through compatible MCP clients. With this server, you can:
Capture webcam images: Get the latest picture from your webcam to show objects or your environment
Take screenshots: Capture your current screen or window (automatically resized)
View live webcam feed: Access real-time video from your webcam
Freeze and analyze images: Pause the current webcam view for detailed examination
Send sampling requests: Submit images with questions (e.g., "What am I holding?") to clients that support this feature
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Webcam Servercapture what's on my desk right now"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
⭐⭐ mcp-webcam 0.2.0 - the 50 Star Update ⭐⭐
In celebration of getting 52 GitHub stars, mcp-webcam 0.2.0 is here! Now supports streamable-http!! No installation required! - try it now at https://webcam.fast-agent.ai/. You can specify your own UserID by adding ?user=<YOUR_USER_ID> after the URL. Note this shared instance is for fun, not security - see below for instructions how to run your own copy locally.
In streamable-http mode multiple clients can connect simultaneously, and you can choose which is used for Sampling.
If we get to 100 stars I'll add another feature 😊.
Multi-user Mode
When run in Streaming mode, if you set an MCP_HOST environment variable the host name is used as a prefix in URL construction, and 5 character UserIDs are automatically generated when the User lands on the webpage.
Related MCP server: Webcam MCP
mcp-webcam
MCP Server that provides access to your WebCam. Provides capture and screenshot tools to take an image from the Webcam, or take a screenshot. The current image is also available as a Resource.
MCP Sampling
mcp-webcam supports "sampling"! Press the "Sample" button to send a sampling request to the Client along with your entered message.
Claude Desktop does not currently support Sampling. If you want a Client that can handle multi-modal sampling request, tryhttps://github.com/evalstate/fast-agent/ or VSCode (more details below).
Installation and Running
NPX
Install a recent version of NodeJS for your platform. The NPM package is @llmindset/mcp-webcam.
To start in STDIO mode: npx @llmindset/mcp-webcam. This starts the mcp-webcam UI on port 3333. Point your browser at http://localhost:3333 to get started.
To change the port: npx @llmindset/mcp-webcam 9999. This starts mcp-webcam the UI on port 9999.
For Streaming HTTP mode: npx @llmindset/mcp-webcam --streaming. This will make the UI available at http://localhost:3333 and the MCP Server available at http://localhost:3333/mcp.
Docker
You can run mcp-webcam using Docker. By default, it starts in streaming mode:
docker run -p 3333:3333 ghcr.io/evalstate/mcp-webcam:latestEnvironment Variables
MCP_TRANSPORT_MODE- Set tostdiofor STDIO mode, defaults tostreamingPORT- The port to run on (default:3333)BIND_HOST- Network interface to bind the server to (default:localhost)MCP_HOST- Public-facing URL for user instructions and MCP client connections (default:http://localhost:3333)
Examples
# STDIO mode
docker run -p 3333:3333 -e MCP_TRANSPORT_MODE=stdio ghcr.io/evalstate/mcp-webcam:latest
# Custom port
docker run -p 8080:8080 -e PORT=8080 ghcr.io/evalstate/mcp-webcam:latest
# For cloud deployments with custom domain (e.g., Hugging Face Spaces)
docker run -p 3333:3333 -e MCP_HOST=https://evalstate-mcp-webcam.hf.space ghcr.io/evalstate/mcp-webcam:latest
# Complete cloud deployment example
docker run -p 3333:3333 -e MCP_HOST=https://your-domain.com ghcr.io/evalstate/mcp-webcam:latestClients
If you want a Client that supports sampling try:
fast-agent
Start the mcp-webcam in streaming mode, install uv and connect with:
uvx fast-agent-mcp go --url http://localhost:3333/mcp
fast-agent currently uses Haiku as its default model, so set an ANTHROPIC_API_KEY. If you want to use a different model, you can add --model on the command line. More instructions for installation and configuration are available here: https://fast-agent.ai/models/.
To start the server in STDIO mode, add the following to your fastagent.config.yaml
webcam_local:
command: "npx"
args: ["@llmindset/mcp-webcam"]VSCode
VSCode versions 1.101.0 and above support MCP Sampling. Simply start mcp-webcam in streaming mode, and add http://localhost:3333/mcp as an MCP Server to get started.
Claude Desktop
Claude Desktop does NOT support Sampling. To run mcp-webcam from Claude Desktop, add the following to the mcpServers section of your claude_desktop_config.json file:
"webcam": {
"command": "npx",
"args": [
"-y",
"@llmindset/mcp-webcam"
]
}Start Claude Desktop, and connect to http://localhost:3333. You can then ask Claude to get the latest picture from my webcam, or Claude, take a look at what I'm holding or what colour top am i wearing?. You can "freeze" the current image and that will be returned to Claude rather than a live capture.
You can ask for Screenshots - navigate to the browser so that you can guide the capture area when the request comes in. Screenshots are automatically resized to be manageable for Claude (useful if you have a 4K Screen). The button is there to allow testing of your platform specific Screenshot UX - it doesn't do anything other than prepare you for a Claude intiated request. NB this does not not work on Safari as it requires human initiation.
Other notes
That's it really.
This MCP Server was built to demonstrate exposing a User Interface on an MCP Server, and serving live resources back to Claude Desktop.
This project might prove useful if you want to build a local, interactive MCP Server.
Thanks to https://github.com/tadasant for help with testing and setup.
Please read the article at https://llmindset.co.uk/posts/2025/01/resouce-handling-mcp for more details about handling files and resources in LLM / MCP Chat Applications, and why you might want to do this.
Available Tools
2 toolscaptureARead-only
Gets the latest picture from the webcam. You can use this if the human asks questions about their immediate environment, if you want to see the human or to examine an object they may be referring to or showing you.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true and openWorldHint=true, indicating safe, non-destructive operation with potential for varied outcomes. The description adds valuable context by specifying it captures from 'the webcam' and returns 'the latest picture,' clarifying the source and immediacy of the data, which goes beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by usage guidelines in a clear, efficient manner. Every sentence adds value without redundancy, making it appropriately sized and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no output schema), the description is complete enough for effective use. It covers purpose, usage guidelines, and behavioral context. The absence of an output schema is mitigated by the description's clarity on what is returned ('the latest picture'), though more detail on output format could enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately does not discuss parameters, maintaining focus on tool functionality. A baseline of 4 is applied since there are no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Gets the latest picture from the webcam') and resource ('webcam'), distinguishing it from the sibling tool 'screenshot' which likely captures screen content rather than camera input. The verb 'Gets' is precise and the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides when-to-use guidance with concrete examples: 'if the human asks questions about their immediate environment,' 'if you want to see the human,' or 'to examine an object they may be referring to or showing you.' This gives clear context for selecting this tool over alternatives like 'screenshot'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotBRead-only
Gets a screenshot of the current screen or window
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and openWorldHint=true, so the agent knows this is a safe, non-destructive operation with open-world assumptions. The description adds minimal behavioral context beyond this, such as specifying it captures the 'current screen or window', but doesn't detail aspects like format, size, or potential limitations. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that efficiently conveys the core functionality without any wasted words. It is front-loaded with the essential information, making it easy for an agent to parse and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema) and rich annotations (readOnlyHint, openWorldHint), the description is adequate but minimal. It covers the basic purpose but lacks details on output format or behavioral nuances that could aid the agent, such as whether it returns an image file or data. It meets minimum viability but has gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing instead on the tool's action. A baseline of 4 is applied since it avoids redundancy and adds value by explaining the tool's purpose without unnecessary details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Gets') and resource ('screenshot of the current screen or window'), making it immediately understandable. However, it doesn't explicitly differentiate from the sibling tool 'capture', which might have overlapping functionality, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling tool 'capture', nor does it mention any prerequisites, context, or exclusions. It merely states what the tool does without offering usage instructions, leaving the agent to infer when it's appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
capture - First observed
screenshot
TDQS
The tools have overlapping purposes—both capture visual data from the user's environment—with 'capture' targeting the webcam and 'screenshot' targeting the screen, but descriptions could lead to confusion as 'capture' mentions examining objects the human shows, which might overlap with screen content. The boundaries are somewhat unclear, especially for agents interpreting use cases.
Tool names follow a consistent verb-based pattern ('capture' and 'screenshot'), both being single words describing the action. There are no deviations in style or casing, making them readable and predictable, though 'screenshot' is more specific than 'capture' in terms of naming convention.
With only 2 tools, the server feels under-scoped for a webcam domain, as it lacks operations like video capture, settings adjustment, or multi-camera support. This minimal set may limit agent functionality, making it borderline too few for comprehensive visual input handling.
The tool surface is significantly incomplete for a webcam server; it covers basic image capture from webcam and screen but misses essential operations such as starting/stopping video, configuring camera settings, or handling multiple inputs. This creates gaps that could lead to agent failures in more complex visual tasks.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Generate AI images and videos from any compatible MCP client.
Holiday photo MCP server: list and fetch personal holiday photos inline in Claude chat.
Related MCP Servers
- -licenseNot gradedqualityNot gradedmaintenanceProvides real-time screen capture streaming in base64 format via WebSocket connections, enabling Claude to view and analyze user screens through natural language requests.1-
- AlicenseNot gradedqualityCmaintenanceAn MCP server that provides LLM agents with direct access to webcam hardware for capturing high-resolution photos and recording video sequences. It enables autonomous agents to monitor environments and interact with the physical world through standard Model Context Protocol tools.3MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with local camera devices to capture and process images. It allows LLMs to access video devices with configurable settings such as resolution, orientation, and image format.1MIT
- FlicenseAqualityDmaintenanceEnables LLMs to capture screenshots and screen recordings through MCP with chunked session-based transfers for reliable image consumption. Supports multi-monitor selection, timeline capture, and compatibility with both vision and non-vision language models.111-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/evalstate/mcp-webcam'
If you have feedback or need assistance with the MCP directory API, please join our Discord server