Skip to main content
Glama

images-handler

给只支持文本的模型(如 DeepSeek)补上"看图"能力的标准 MCP 服务。

DeepSeek 不能直接识别图片,但本服务通过 Cursor TypeScript SDK 在本机跑一个 Cursor agent(默认 composer-2,可换 Claude/GPT 视觉模型),把图片理解成文本描述返回。DeepSeek 调用工具拿到文字结果,就等于"能看图"了。

本服务只做图片识别:agent 始终以纯文本模式运行,不执行任何 shell/文件工具。

前置条件

  • Node.js ≥ 22.13

  • Cursor 凭据(二选一):

    • 设置环境变量 CURSOR_API_KEY,或

    • 已用 Cursor.auth.login() 登录过 Cursor 账号(SDK 自动读取存储的凭据)

Related MCP server: DeepSeek Eyes

安装与运行

npm install
npm start            # 开发运行(stdio),等价 npx tsx src/index.ts
# 或构建后运行
npm run build && node dist/index.js

环境变量

变量

默认

说明

CURSOR_API_KEY

Cursor API key,缺省时回退登录态

CURSOR_MODEL

composer-2

视觉模型 id

CURSOR_AGENT_TIMEOUT_MS

600000

单次识别调用超时(毫秒)

接入客户端

Claude Code

claude mcp add image-recognition -e CURSOR_API_KEY="${CURSOR_API_KEY}" -- npx tsx D:/path/to/images-handler/src/index.ts

Cursor

.cursor/mcp.json:

{
  "mcpServers": {
    "image-recognition": {
      "command": "npx",
      "args": ["tsx", "D:/path/to/images-handler/src/index.ts"],
      "env": {
        "CURSOR_API_KEY": "${CURSOR_API_KEY}"
      }
    }
  }
}

其他标准 MCP 客户端

stdio 传输,按标准协议配置启动命令即可(记得通过 env 传入 CURSOR_API_KEY)。

工具:recognize_image

参数

类型

必填

说明

images

string 或 string[]

图片 data URI(data:image/png;base64,...)、http(s) URL 或本地图片文件路径(如 D:/photos/a.png)

instruction

string

想针对图片问什么,缺省为"请详细描述这张图片的内容、画面元素和任何可见文字。"

model

string

覆盖视觉模型(默认 composer-2)

示例

{
  "images": ["data:image/png;base64,iVBORw0KGgo..."]
}

带自定义指令:

{
  "images": ["data:image/png;base64,iVBORw0KGgo..."],
  "instruction": "识别图中的文字并翻译成中文"
}

传本地文件路径(在服务所在机器上读取):

{
  "images": ["D:/photos/screenshot.png"]
}

说明与限制

  • 每次工具调用都会新建一个独立 Cursor agent(调用间不共享会话历史),用完即关闭。

  • 服务只做图片识别:agent 恒为纯文本模式(tools: []),不执行 shell/文件工具,除传入的图片外不会读取或访问任何本地内容。

  • 传本地路径时,文件在服务所在机器上读取,并按扩展名识别为图片;非图片扩展名会被拒绝。请仅传入你自己信任的图片路径。

Available Tools

1 tool
recognize_imageA

Recognize and analyze the given image(s) with a vision model and return a text description. Use this to give text-only LLMs like DeepSeek image understanding.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoVision model id (default: composer-2).
imagesYesImage data URI (data:image/png;base64,...), http(s) URL, or local image file path to recognize.
instructionNoWhat to ask about the image. Default: 请详细描述这张图片的内容、画面元素和任何可见文字。

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the use of a vision model and the return type (text description), but it does not mention potential side effects like data transmission to an external service, latency, or failure modes. The disclosure is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loads the primary purpose, and contains no superfluous words. Every phrase contributes to understanding what the tool does and why to use it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters and no output schema, the description sufficiently covers the purpose and use case. The schema handles parameter details and defaults, so the description does not need to elaborate further. Minor gaps like clarifying multi-image behavior are already inferable from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all three parameters (model, images, instruction). The description adds no parameter-specific details beyond what the schema already provides, so it stays at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs 'recognize and analyze' and clearly identifies the resource as 'image(s)' and output as 'text description'. It also adds a concrete use case ('give text-only LLMs like DeepSeek image understanding'), making the tool's intent unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides an explicit when-to-use context: 'Use this to give text-only LLMs like DeepSeek image understanding.' However, it does not mention when not to use it or any alternative tools, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 1 tool updatev0.1.0
    • First observedrecognize_image

TDQS

A3.9/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The single tool has a clear and distinct purpose.

Naming Consistency5/5

The tool name 'recognize_image' follows a clear verb_noun convention. Although there are no other tools to compare, the naming is internally consistent and predictable.

Tool Count3/5

The server has only 1 tool, which feels thin for an 'images-handler' name that implies a broader scope. However, the single tool is non-trivial and provides meaningful functionality, so it is borderline rather than severely inadequate.

Completeness2/5

The server only offers image recognition/analysis, missing any other image handling operations like editing, conversion, or resizing. Given the 'images-handler' naming, this is a significant gap that limits the server's usefulness for general image tasks.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.
    1,060
    10
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that grants image recognition to text-only models like DeepSeek by forwarding images to vision models and returning text descriptions. Supports clipboard, pasted session images, and batch folder image recognition.
    5
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server that gives text-only LLMs like DeepSeek vision capabilities by converting images to text via vision APIs, enabling image description, OCR, and generation in MCP clients.
    1
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/yuzhaoyang001/my_mcps-image-handler'

If you have feedback or need assistance with the MCP directory API, please join our Discord server