image-analysis-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@image-analysis-mcpextract text from ~/screenshots/page.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
image-analysis-mcp
FastMCP server for image analysis — OCR, metadata, and EXIF extraction. Part of the Palimpsest intelligence toolkit.
Why
LLMs can't see images natively. This server fills the gap — extract text from screenshots, photos, and diagrams via OCR, pull EXIF camera data and GPS coordinates, and get full image metadata (format, dimensions, DPI, color space). All results are returned as structured JSON for easy downstream processing.
Related MCP server: mcp-ocrspace
Architecture
Pluggable OCR backends with graceful degradation:
base install (
pip install image-analysis-mcp) — metadata / EXIF only (Pillow + exifread, ~5 MB)with OCR (
pip install image-analysis-mcp[ocr]) — adds rapidocr-onnxruntime (~250 MB)with Tesseract (
pip install image-analysis-mcp[tesseract]) — adds pytesseract (needs systemtesseract-ocr)
The server tries backends in priority order: rapidocr → tesseract → none.
Tools
Tool | Description |
| OCR text from an image — returns text blocks with confidence scores and bounding boxes |
| Full metadata: file info, image dimensions, EXIF tags, GPS coordinates |
| Convenience wrapper — OCR + metadata combined in one response |
Installation
git clone https://github.com/palimpsest-labs/image-analysis-mcp
cd image-analysis-mcp
python3 -m venv .venv
source .venv/bin/activate
# Minimal (metadata only)
pip install -e .
# With OCR support
pip install -e ".[ocr]"
# With Tesseract support (requires system tesseract-ocr)
pip install -e ".[tesseract]"Usage
from image_analysis_mcp.server import extract_text, image_metadata, ocr_image
# Get everything at once
result = extract_text("~/screenshots/page.png")
# Or separate calls
meta = image_metadata("~/screenshots/page.png")
text = ocr_image("~/screenshots/page.png")Security
Paths must be absolute and resolve to a location under the user's home directory. Path traversal (..) and paths starting with / are rejected. Symlinks are resolved before checking home-directory containment.
License
MIT
Available Tools
3 toolsextract_textA
Extract all text and metadata from an image.
Convenience wrapper that combines OCR text extraction with full image metadata in a single JSON response.
Args: image_path: Absolute path to the image file (must be under home directory)
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It usefully reveals that the tool returns a 'single JSON response' and imposes a security constraint ('must be under home directory'). However, it does not mention error handling, supported file types, or explicitly state that it is a read-only operation, leaving some behavioral aspects undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the primary purpose. It uses three concise sentences with no unnecessary fluff, and the Args section clearly explains the parameter without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, the description covers the essential invocation details: what it does, the path requirement, and the combined output. It does not specify image format limitations or failure modes, but these are not critical for a basic extraction tool. Overall, it is adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides no description for the single parameter (coverage 0%), so the description fully compensates by explaining that image_path must be an absolute path under the home directory. This adds crucial semantic meaning beyond the schema's bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Extract all text and metadata from an image' with a specific verb and resource. It also distinguishes itself from siblings by calling itself a 'convenience wrapper' that combines OCR text extraction with image metadata, making it distinct from the standalone ocr_image and image_metadata tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by labeling the tool a 'convenience wrapper' that combines OCR and metadata into one response, suggesting it should be used when both are needed. However, it does not explicitly state when not to use it or directly reference the sibling alternatives, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image_metadataA
Extract full metadata from an image.
Returns file info (size, timestamps, SHA-256), image properties (format, dimensions, DPI, colour space), and EXIF data including camera make/model, datetime, and GPS coordinates.
Args: path: Absolute path to the image file (must be under home directory)
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It clearly indicates a read-only operation via 'extract' and details what the tool returns, plus the path constraint that the file must be under home directory. However, it does not mention error handling, file type limitations, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, opening with a direct purpose statement followed by a focused list of return categories. The explicit Args section adds necessary detail without redundancy, and every sentence contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description provides sufficient context: it states the purpose, lists the returned metadata types, and explains the path constraint. Despite lacking explicit usage guidance, the core operation is fully specified for the agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a type-less 'path' property with no description. The description compensates fully by defining 'path' as 'Absolute path to the image file' and adding the additional constraint that it must be under the home directory, giving the agent complete parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Extract full metadata from an image,' which clearly states the verb and resource. It then enumerates specific metadata categories (file info, image properties, EXIF), distinguishing it from sibling tools like ocr_image and extract_text that handle text extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly guide when to use this tool versus alternatives. While it is clear the tool is for metadata extraction, it never mentions the text-focused siblings or provides when/when-not guidance, leaving usage to be implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_imageA
OCR text from an image.
Extracts text using the best available OCR backend. Returns text blocks with confidence scores and bounding boxes.
Args: image_path: Absolute path to the image file (must be under home directory)
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return format (text blocks with confidence scores and bounding boxes), which is useful, but it does not mention potential errors, permissions, side effects, or backend behavior with edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose ('OCR text from an image'), followed by a brief explanation of the output. Every sentence is informative, with no wasteful repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter, and the description covers the invocation and output format. The presence of an output schema reduces the need to describe return values in detail. However, it lacks usage guidance and constraints like supported image formats or image size limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful details for 'image_path' beyond the schema, specifying it must be an absolute path and under the home directory. Since schema coverage is 0%, this description compensates well for the single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb 'OCR' and resource 'image', clearly defining the tool's function. It also mentions the output type (text blocks with confidence scores and bounding boxes), but it does not explicitly distinguish from the sibling 'extract_text', which could overlap in functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide guidance on when to use this tool versus alternatives like 'extract_text' or 'image_metadata'. It only implies usage for image files but lacks explicit context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v0.1.0- First observed
extract_text - First observed
image_metadata - First observed
ocr_image
TDQS
ocr_image and extract_text both handle text extraction, causing slight overlap, but extract_text explicitly combines OCR with metadata, making its broader purpose clear. image_metadata is distinct.
Names mix conventions: ocr_image (verb_noun), image_metadata (noun_noun), and extract_text (verb_noun). The inconsistency is noticeable but still readable.
Three tools is well-scoped for a focused image analysis server, though extract_text is somewhat redundant as a convenience wrapper.
The server covers OCR and metadata extraction, including a combined wrapper, but lacks broader image analysis features such as classification or object detection, which the server name might imply.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
OCR.space MCP — wraps the OCR.space API (ocr.space) for image/PDF → text OCR.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server for web extraction and rendering via AceDataCloud WebExtrator
MCP server for detecting and redacting PII (Personally Identifiable Information) in PDF documents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that provides OCR capabilities using the EasyOCR library, supporting over 80 languages and GPU acceleration. It enables processing images from base64 strings, local files, or URLs with options for text-only or detailed coordinate and confidence output.2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceOCR.space MCP server that enables image and PDF text extraction via the OCR.space API.17MIT
- AlicenseAqualityDmaintenanceHigh-performance OCR MCP server supporting multiple input modes (path, base64, URL, upload), batch processing, and output formats like plain, JSON, and Markdown.42MIT
- AlicenseAqualityDmaintenanceMCP server providing image analysis tools for AI agents, including metadata extraction, favicon discovery, and placeholder generation.551MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/jmars/image-analysis-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server