Skip to main content
Glama
jmars
by jmars

image-analysis-mcp

FastMCP server for image analysis — OCR, metadata, and EXIF extraction. Part of the Palimpsest intelligence toolkit.

Why

LLMs can't see images natively. This server fills the gap — extract text from screenshots, photos, and diagrams via OCR, pull EXIF camera data and GPS coordinates, and get full image metadata (format, dimensions, DPI, color space). All results are returned as structured JSON for easy downstream processing.

Related MCP server: mcp-ocrspace

Architecture

Pluggable OCR backends with graceful degradation:

  • base install (pip install image-analysis-mcp) — metadata / EXIF only (Pillow + exifread, ~5 MB)

  • with OCR (pip install image-analysis-mcp[ocr]) — adds rapidocr-onnxruntime (~250 MB)

  • with Tesseract (pip install image-analysis-mcp[tesseract]) — adds pytesseract (needs system tesseract-ocr)

The server tries backends in priority order: rapidocr → tesseract → none.

Tools

Tool

Description

ocr_image(image_path)

OCR text from an image — returns text blocks with confidence scores and bounding boxes

image_metadata(path)

Full metadata: file info, image dimensions, EXIF tags, GPS coordinates

extract_text(image_path)

Convenience wrapper — OCR + metadata combined in one response

Installation

git clone https://github.com/palimpsest-labs/image-analysis-mcp
cd image-analysis-mcp
python3 -m venv .venv
source .venv/bin/activate

# Minimal (metadata only)
pip install -e .

# With OCR support
pip install -e ".[ocr]"

# With Tesseract support (requires system tesseract-ocr)
pip install -e ".[tesseract]"

Usage

from image_analysis_mcp.server import extract_text, image_metadata, ocr_image

# Get everything at once
result = extract_text("~/screenshots/page.png")

# Or separate calls
meta = image_metadata("~/screenshots/page.png")
text = ocr_image("~/screenshots/page.png")

Security

Paths must be absolute and resolve to a location under the user's home directory. Path traversal (..) and paths starting with / are rejected. Symlinks are resolved before checking home-directory containment.

License

MIT

Available Tools

3 tools
extract_textA

Extract all text and metadata from an image.

Convenience wrapper that combines OCR text extraction with full image metadata in a single JSON response.

Args: image_path: Absolute path to the image file (must be under home directory)

ParametersJSON Schema
NameRequiredDescriptionDefault
image_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It usefully reveals that the tool returns a 'single JSON response' and imposes a security constraint ('must be under home directory'). However, it does not mention error handling, supported file types, or explicitly state that it is a read-only operation, leaving some behavioral aspects undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the primary purpose. It uses three concise sentences with no unnecessary fluff, and the Args section clearly explains the parameter without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description covers the essential invocation details: what it does, the path requirement, and the combined output. It does not specify image format limitations or failure modes, but these are not critical for a basic extraction tool. Overall, it is adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides no description for the single parameter (coverage 0%), so the description fully compensates by explaining that image_path must be an absolute path under the home directory. This adds crucial semantic meaning beyond the schema's bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Extract all text and metadata from an image' with a specific verb and resource. It also distinguishes itself from siblings by calling itself a 'convenience wrapper' that combines OCR text extraction with image metadata, making it distinct from the standalone ocr_image and image_metadata tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by labeling the tool a 'convenience wrapper' that combines OCR and metadata into one response, suggesting it should be used when both are needed. However, it does not explicitly state when not to use it or directly reference the sibling alternatives, so it stops short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image_metadataA

Extract full metadata from an image.

Returns file info (size, timestamps, SHA-256), image properties (format, dimensions, DPI, colour space), and EXIF data including camera make/model, datetime, and GPS coordinates.

Args: path: Absolute path to the image file (must be under home directory)

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It clearly indicates a read-only operation via 'extract' and details what the tool returns, plus the path constraint that the file must be under home directory. However, it does not mention error handling, file type limitations, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, opening with a direct purpose statement followed by a focused list of return categories. The explicit Args section adds necessary detail without redundancy, and every sentence contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description provides sufficient context: it states the purpose, lists the returned metadata types, and explains the path constraint. Despite lacking explicit usage guidance, the core operation is fully specified for the agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only a type-less 'path' property with no description. The description compensates fully by defining 'path' as 'Absolute path to the image file' and adding the additional constraint that it must be under the home directory, giving the agent complete parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Extract full metadata from an image,' which clearly states the verb and resource. It then enumerates specific metadata categories (file info, image properties, EXIF), distinguishing it from sibling tools like ocr_image and extract_text that handle text extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly guide when to use this tool versus alternatives. While it is clear the tool is for metadata extraction, it never mentions the text-focused siblings or provides when/when-not guidance, leaving usage to be implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ocr_imageA

OCR text from an image.

Extracts text using the best available OCR backend. Returns text blocks with confidence scores and bounding boxes.

Args: image_path: Absolute path to the image file (must be under home directory)

ParametersJSON Schema
NameRequiredDescriptionDefault
image_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return format (text blocks with confidence scores and bounding boxes), which is useful, but it does not mention potential errors, permissions, side effects, or backend behavior with edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core purpose ('OCR text from an image'), followed by a brief explanation of the output. Every sentence is informative, with no wasteful repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter, and the description covers the invocation and output format. The presence of an output schema reduces the need to describe return values in detail. However, it lacks usage guidance and constraints like supported image formats or image size limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful details for 'image_path' beyond the schema, specifying it must be an absolute path and under the home directory. Since schema coverage is 0%, this description compensates well for the single parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'OCR' and resource 'image', clearly defining the tool's function. It also mentions the output type (text blocks with confidence scores and bounding boxes), but it does not explicitly distinguish from the sibling 'extract_text', which could overlap in functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide guidance on when to use this tool versus alternatives like 'extract_text' or 'image_metadata'. It only implies usage for image files but lacks explicit context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.1.0
    • First observedextract_text
    • First observedimage_metadata
    • First observedocr_image

TDQS

A3.9/5.0
Disambiguation4/5

ocr_image and extract_text both handle text extraction, causing slight overlap, but extract_text explicitly combines OCR with metadata, making its broader purpose clear. image_metadata is distinct.

Naming Consistency3/5

Names mix conventions: ocr_image (verb_noun), image_metadata (noun_noun), and extract_text (verb_noun). The inconsistency is noticeable but still readable.

Tool Count5/5

Three tools is well-scoped for a focused image analysis server, though extract_text is somewhat redundant as a convenience wrapper.

Completeness4/5

The server covers OCR and metadata extraction, including a combined wrapper, but lacks broader image analysis features such as classification or object detection, which the server name might imply.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jmars/image-analysis-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server