Gemini Image Generator MCP Server
The Gemini Image Generator MCP Server enables AI assistants to generate and transform images using Google's Gemini model via the MCP protocol.
Generate images from text prompts using
generate_image_from_textTransform existing images using base64-encoded data or from file paths
Automatic features including intelligent filename generation, language translation for prompts, and high-resolution output
Local storage saves generated/transformed images to a configurable directory
Direct access to both raw image data and file paths for AI assistants
Supports environment variable configuration through .env files for storing API keys and output path settings.
Enables text-to-image generation and image transformation using Google's Gemini AI model, supporting high-resolution image creation from text prompts and modification of existing images based on textual descriptions.
Includes specific configuration paths for macOS users to set up the MCP server with Claude Desktop.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini Image Generator MCP Servergenerate a photorealistic sunset over mountains"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini Image Generator MCP Server
Generate high-quality images from text prompts using Google's Gemini model through the MCP protocol.
Overview
This MCP server allows any AI assistant to generate images using Google's Gemini AI model. The server handles prompt engineering, text-to-image conversion, filename generation, and local image storage, making it easy to create and manage AI-generated images through any MCP client.
Related MCP server: Gemini Image Gen MCP Server
Features
Text-to-image generation using Gemini 2.0 Flash
Image-to-image transformation based on text prompts
Support for both file-based and base64-encoded images
Automatic intelligent filename generation based on prompts
Automatic translation of non-English prompts
Local image storage with configurable output path
Strict text exclusion from generated images
High-resolution image output
Direct access to both image data and file path
Available MCP Tools
The server provides the following MCP tools for AI assistants:
1. generate_image_from_text
Creates a new image from a text prompt description.
generate_image_from_text(prompt: str) -> Tuple[bytes, str]Parameters:
prompt: Text description of the image you want to generate
Returns:
A tuple containing:
Raw image data (bytes)
Path to the saved image file (str)
This dual return format allows AI assistants to either work with the image data directly or reference the saved file path.
Examples:
"Generate an image of a sunset over mountains"
"Create a photorealistic flying pig in a sci-fi city"
Example Output
This image was generated using the prompt:
"Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?"
A 3D rendered pig with wings and a top hat flying over a futuristic sci-fi city filled with greenery
Known Issues
When using this MCP server with Claude Desktop Host:
Performance Issues: Using
transform_image_from_encodedmay take significantly longer to process compared to other methods. This is due to the overhead of transferring large base64-encoded image data through the MCP protocol.Path Resolution Problems: There may be issues with correctly resolving image paths when using Claude Desktop Host. The host application might not properly interpret the returned file paths, making it difficult to access the generated images.
For the best experience, consider using alternative MCP clients or the transform_image_from_file method when possible.
2. transform_image_from_encoded
Transforms an existing image based on a text prompt using base64-encoded image data.
transform_image_from_encoded(encoded_image: str, prompt: str) -> Tuple[bytes, str]Parameters:
encoded_image: Base64 encoded image data with format header (must be in format: "data:image/[format];base64,[data]")prompt: Text description of how you want to transform the image
Returns:
A tuple containing:
Raw transformed image data (bytes)
Path to the saved transformed image file (str)
Example:
"Add snow to this landscape"
"Change the background to a beach"
3. transform_image_from_file
Transforms an existing image file based on a text prompt.
transform_image_from_file(image_file_path: str, prompt: str) -> Tuple[bytes, str]Parameters:
image_file_path: Path to the image file to be transformedprompt: Text description of how you want to transform the image
Returns:
A tuple containing:
Raw transformed image data (bytes)
Path to the saved transformed image file (str)
Examples:
"Add a llama next to the person in this image"
"Make this daytime scene look like night time"
Example Transformation
Using the flying pig image created above, we applied a transformation with the following prompt:
"Add a cute baby whale flying alongside the pig"Before:

After:

The original flying pig image with a cute baby whale added flying alongside it
Setup
Prerequisites
Python 3.11+
Google AI API key (Gemini)
MCP host application (Claude Desktop App, Cursor, or other MCP-compatible clients)
Getting a Gemini API Key
Sign in with your Google account
Click "Create API Key"
Copy your new API key for use in the configuration
Note: The API key provides a certain quota of free usage per month. You can check your usage in the Google AI Studio
Installation
Installing via Smithery
To install Gemini Image Generator MCP for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @qhdrl12/mcp-server-gemini-image-gen --client claudeManual Installation
Clone the repository:
git clone https://github.com/your-username/mcp-server-gemini-image-generator.git
cd mcp-server-gemini-image-generatorCreate a virtual environment and install dependencies:
# Using uv (recommended)
uv venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
uv pip install -e .
# Or using regular venv
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
pip install -e .Set up environment variables (choose one method):
Method A: Using .env file (optional)
# Create .env file in the project root
cat > .env << 'EOF'
GEMINI_API_KEY=your-gemini-api-key-here
OUTPUT_IMAGE_PATH=/path/to/save/images
EOFMethod B: Set directly in Claude Desktop config (recommended)
Set environment variables directly in the
claude_desktop_config.json(shown in configuration section below)
Configure Claude Desktop
Add the following to your claude_desktop_config.json:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"gemini-image-generator": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/mcp-server-gemini-image-generator",
"run",
"mcp-server-gemini-image-generator"
],
"env": {
"GEMINI_API_KEY": "your-actual-gemini-api-key-here",
"OUTPUT_IMAGE_PATH": "/absolute/path/to/your/images/directory"
}
}
}
}Important Configuration Notes:
Replace paths with your actual paths:
Change
/absolute/path/to/mcp-server-gemini-image-generatorto the actual location where you cloned this repositoryChange
/absolute/path/to/your/images/directoryto where you want generated images to be saved
Environment Variables:
Replace
your-actual-gemini-api-key-herewith your real Gemini API key from Google AI StudioUse absolute paths for
OUTPUT_IMAGE_PATHto ensure images are saved correctly
Example with real paths:
{
"mcpServers": {
"gemini-image-generator": {
"command": "uv",
"args": [
"--directory",
"/Users/username/Projects/mcp-server-gemini-image-generator",
"run",
"mcp-server-gemini-image-generator"
],
"env": {
"GEMINI_API_KEY": "GEMINI_API_KEY",
"OUTPUT_IMAGE_PATH": "OUTPUT_IMAGE_PATH"
}
}
}
}Usage
Once installed and configured, you can ask Claude to generate or transform images using prompts like:
Generating New Images
"Generate an image of a sunset over mountains"
"Create an illustration of a futuristic cityscape"
"Make a picture of a cat wearing sunglasses"
Transforming Existing Images
"Transform this image by adding snow to the scene"
"Edit this photo to make it look like it was taken at night"
"Add a dragon flying in the background of this picture"
The generated/transformed images will be saved to your configured output path and displayed in Claude. With the updated return types, AI assistants can also work directly with the image data without needing to access the saved files.
Testing
You can test the application by running the FastMCP development server:
fastmcp dev server.pyThis command starts a local development server and makes the MCP Inspector available at http://localhost:5173/. The MCP Inspector provides a convenient web interface where you can directly test the image generation tool without needing to use Claude or another MCP client. You can enter text prompts, execute the tool, and see the results immediately, which is helpful for development and debugging.
License
MIT License
Available Tools
3 toolsgenerate_image_from_textA
Generate an image based on the given text prompt using Google's Gemini model.
Args:
prompt: User's text prompt describing the desired image to generate
Returns:
Path to the generated image file using Gemini's image generation capabilities
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the model and return type (path to image file) but lacks critical details such as rate limits, authentication requirements, image format, size, quality, or error handling. This is insufficient for a generative AI tool with potential costs and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, starting with the core functionality. The structured sections (Args, Returns) enhance readability, though the second sentence could be more integrated to avoid slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of image generation, no annotations, and no output schema, the description is incomplete. It lacks details on behavioral traits (e.g., costs, latency), output specifics (e.g., file format, resolution), and error cases, leaving significant gaps for an AI agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, but the description compensates by explaining the single parameter ('prompt') as 'User's text prompt describing the desired image to generate.' This adds meaningful context beyond the schema's basic type information, clarifying the parameter's role in the generation process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate an image') and resource ('based on the given text prompt'), using Google's Gemini model. It distinguishes from sibling tools like 'transform_image_from_encoded' and 'transform_image_from_file' by specifying text-based generation rather than transformation from existing images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for text-to-image generation but does not explicitly state when to use this tool versus alternatives. It mentions the model (Gemini) but provides no guidance on prerequisites, limitations, or scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_encodedA
Transform an existing image based on the given text prompt using Google's Gemini model.
Args:
encoded_image: Base64 encoded image data with header. Must be in format:
"data:image/[format];base64,[data]"
Where [format] can be: png, jpeg, jpg, gif, webp, etc.
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| encoded_image | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the tool uses Google's Gemini model and that it saves the transformed image on the server, which are useful behavioral traits. However, it doesn't mention rate limits, authentication requirements, file size limits, or potential side effects of the transformation process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear opening sentence stating the purpose, followed by well-organized sections for Args and Returns. Every sentence earns its place by providing essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no annotations and no output schema, the description provides good coverage of purpose, parameters, and basic behavior. It explains what the tool does, how to format inputs, and what to expect as output. The main gap is lack of information about error conditions, performance characteristics, or more detailed behavioral constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by providing detailed semantics for both parameters. It specifies the exact format required for encoded_image (including header format and supported image types) and explains what the prompt parameter should contain. This adds significant value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('an existing image'), and distinguishes it from siblings by specifying it uses encoded image data rather than text or file inputs. The mention of Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool (transforming existing images with encoded data) and implicitly distinguishes it from siblings (generate_image_from_text for text-to-image, transform_image_from_file for file-based transformation). However, it doesn't explicitly state when NOT to use this tool or mention specific prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transform_image_from_fileA
Transform an existing image file based on the given text prompt using Google's Gemini model.
Args:
image_file_path: Path to the image file to be transformed
prompt: Text prompt describing the desired transformation or modifications
Returns:
Path to the transformed image file saved on the server
| Name | Required | Description | Default |
|---|---|---|---|
| image_file_path | Yes | ||
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the tool saves the transformed file on the server, which is useful behavioral context. However, it lacks critical details like required permissions, file format limitations, transformation scope, error handling, or whether the operation is reversible/destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear purpose statement followed by labeled sections for Args and Returns. Every sentence adds value without redundancy, and information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 2 parameters, the description covers purpose and parameters adequately. However, for a transformation tool with potential complexity (image processing via Gemini), it lacks details about output format, file location specifics, or error cases, leaving gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining both parameters: 'image_file_path' as 'Path to the image file to be transformed' and 'prompt' as 'Text prompt describing the desired transformation or modifications'. This adds essential meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verb ('Transform') and resource ('existing image file'), and distinguishes it from siblings by specifying it works from a file path rather than text or encoded input. The mention of using Google's Gemini model adds technical specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying it transforms 'an existing image file' and uses a 'text prompt', which differentiates it from 'generate_image_from_text' (creates new images) and 'transform_image_from_encoded' (uses encoded input). However, it doesn't explicitly state when to choose this tool over alternatives or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
- First observed
generate_image_from_text - First observed
transform_image_from_encoded - First observed
transform_image_from_file
TDQS
The three tools have clearly distinct purposes: generate_image_from_text creates new images from text prompts, while transform_image_from_encoded and transform_image_from_file both transform existing images but differ in input format (base64 encoded vs. file path). The descriptions make these distinctions explicit, eliminating any potential confusion between generation and transformation operations.
All tools follow a consistent verb_noun_from_source naming pattern: generate_image_from_text, transform_image_from_encoded, and transform_image_from_file. This pattern clearly indicates the action (generate/transform), the target (image), and the input source (text/encoded/file), creating a predictable and readable naming convention throughout the toolset.
Three tools is a reasonable count for an image generation server, covering the core operations of generating new images and transforming existing ones. However, the scope feels slightly thin as there are no complementary tools for managing generated images (like listing, deleting, or retrieving metadata), which might limit agent workflows in production scenarios.
The server covers basic image generation and transformation operations well, but has notable gaps in image management. There are no tools for listing generated images, deleting files, retrieving image metadata, or batch operations. While the core generative AI functionality is present, the lack of lifecycle management tools creates potential dead ends for agents working with multiple images over time.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate AI images and videos from any compatible MCP client.
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.0 and more.
Related MCP Servers
- FlicenseBqualityDmaintenanceEnables image generation and multi-turn editing sessions using the Gemini API within MCP-compatible environments. Users can create, modify, and configure images through natural language commands, supporting features like aspect ratio adjustments and session-based image transformations.5-
- AlicenseAqualityCmaintenanceEnables AI image generation, editing, and upscaling via Google Gemini and Imagen models, supporting dynamic model switching and multiple MCP-compatible clients.12MIT
- FlicenseNot gradedqualityDmaintenanceProvides image generation capabilities using Google's Gemini 2.0 Flash Preview model through the MCP protocol, enabling AI assistants to generate high-quality images from text prompts.-
- FlicenseNot gradedqualityDmaintenanceEnables AI-powered image generation using Google's Gemini 2.5 Flash Image Preview model, supporting text-to-image and image-to-image generation through the MCP interface.-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/qhdrl12/mcp-server-gemini-image-generator'
If you have feedback or need assistance with the MCP directory API, please join our Discord server