nanobanana-mcp
Provides tools for generating and editing images using Google Gemini models, including support for text-to-image prompts, reference images, and session-based iterative editing.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@nanobanana-mcpGenerate a high-quality photo of a cozy cabin in a snowy forest at sunset"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
nanobanana-mcp
A hardened MCP server for Gemini image generation. Fork of ConechoAI/Nano-Banana-MCP with security fixes, strict TypeScript, and model selection.
Features
3 tools:
generate_image,edit_image,continue_editingModel selection via
NANOBANANA_MODELenv var with whitelist validationSecurity hardened: path traversal protection, file size limits, no plaintext key storage
Strict TypeScript: zero
anytypes, Zod validation on all inputs
Related MCP server: Nano-Banana MCP Server
Quick Start
Claude Code
Add to ~/.claude/settings.json:
{
"mcpServers": {
"nanobanana": {
"command": "npx",
"args": ["tsx", "/path/to/nanobanana-mcp/src/index.ts"],
"env": {
"GEMINI_API_KEY": "your-api-key",
"NANOBANANA_MODEL": "gemini-2.5-flash-image"
}
}
}
}Other MCP Clients
GEMINI_API_KEY=your-key npx tsx src/index.tsThe server communicates over stdio using the MCP protocol.
Tools
generate_image
Generate a new image from a text prompt.
prompt (required): Text describing the image to create (max 10,000 chars)edit_image
Edit an existing image with a text prompt.
imagePath (required): Full file path to the image to edit
prompt (required): Text describing the modifications (max 10,000 chars)
referenceImages (optional): Array of file paths to reference imagescontinue_editing
Continue editing the last generated/edited image in the current session.
prompt (required): Text describing changes to make (max 10,000 chars)
referenceImages (optional): Array of file paths to reference imagesConfiguration
All configuration is via environment variables. No config files are written to disk.
Variable | Required | Description |
| Yes | Google Gemini API key |
| No | Override for |
| No | Model to use (see below) |
Available Models
Model ID | Description |
| Fast generation, good for high-volume use (default) |
| Pro quality, complex prompts, better text rendering |
| Latest model, advanced features |
Output
Generated images are saved to ~/nanobanana-images/ with unique filenames. The tool response includes both the file path and the image data inline.
Security
This fork addresses the following security issues from the original:
Issue | Fix |
API key saved to disk in plaintext | Removed config file persistence entirely |
| Tool removed; keys only via env vars |
Path traversal in |
|
No prompt length validation | Capped at 10,000 chars via Zod |
Hardcoded model |
|
Silent swallowing of reference image errors | Errors now thrown and reported |
|
|
No file size limit on reads | Max 20MB |
Verbose errors leak internal paths | Sanitized error messages |
| Fixed to |
Development
npm install
npm run typecheck # Type check without emitting
npm run dev # Run with tsx (hot reload)
npm run build # Compile to dist/Project Structure
src/
index.ts # MCP server entry point (3 tool handlers)
gemini-client.ts # Gemini API wrapper with model selection
file-handler.ts # Secure file I/O with path validation
types.ts # TypeScript interfaces and Zod schemasLicense
MIT - Based on ConechoAI/Nano-Banana-MCP
Available Tools
3 toolscontinue_editingA
Continue editing the LAST image generated or edited in this session. Automatically uses the previous image without needing a file path. Use for iterative improvements.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text describing changes to make to the last image (max 10,000 chars) | |
| referenceImages | No | Optional array of file paths to reference images |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses key behavioral traits: it operates on session state ('LAST image... in this session'), is stateful ('Automatically uses the previous image'), and is designed for 'iterative improvements'. However, it doesn't mention potential limitations like session duration, error handling, or what happens if no previous image exists, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by operational details and usage context. Every sentence earns its place: the first defines the tool, the second explains automation, and the third provides guidance. It's concise with zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (stateful operation, 2 parameters), no annotations, and no output schema, the description is largely complete. It covers purpose, usage, and key behavior. However, it lacks details on output format or error cases, which could be helpful for an agent. Since there's no output schema, some additional context on returns would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds no specific parameter semantics beyond implying 'prompt' is for 'changes to make' and 'referenceImages' might be optional references. This meets the baseline of 3 since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Continue editing'), the resource ('the LAST image generated or edited in this session'), and distinguishes it from siblings by emphasizing it works on the 'previous image without needing a file path' unlike edit_image which likely requires a file path. The phrase 'iterative improvements' further clarifies its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('Continue editing the LAST image... in this session') and provides clear alternatives by naming sibling tools (edit_image, generate_image) in the context. It specifies 'Automatically uses the previous image' which implies when not to use it (e.g., when starting fresh or editing a different image).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_imageB
Edit an existing image file with a text prompt, optionally using additional reference images. Use this when you have the exact file path of an image to modify.
| Name | Required | Description | Default |
|---|---|---|---|
| imagePath | Yes | Full file path to the image to edit | |
| prompt | Yes | Text describing the modifications to make (max 10,000 chars) | |
| referenceImages | No | Optional array of file paths to reference images (for style transfer, adding elements, etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions the tool edits images with a prompt and optional references, it doesn't disclose critical behavioral traits like whether this is a destructive operation (overwrites the original file?), what permissions are needed, rate limits, output format, or error conditions. For a mutation tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and well-structured: two sentences that efficiently convey purpose and usage guidelines. Every word earns its place with no redundancy or fluff. It's front-loaded with the core functionality followed by the key usage condition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 3 parameters, no annotations, and no output schema, the description is incomplete. It adequately covers purpose and basic usage but lacks crucial behavioral context (destructive nature, permissions, output format) and doesn't compensate for the absence of structured metadata. The agent would be left guessing about important operational aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds minimal value beyond what's in the schema: it mentions 'text prompt' and 'reference images' but doesn't provide additional semantic context like examples, constraints beyond the schema's max length, or how reference images are used. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Edit an existing image file with a text prompt, optionally using additional reference images.' It specifies the verb ('edit'), resource ('existing image file'), and key mechanisms (text prompt, reference images). However, it doesn't explicitly distinguish this from sibling tools like 'continue_editing' or 'generate_image' beyond mentioning file path requirements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage guidance: 'Use this when you have the exact file path of an image to modify.' This gives a specific when-to-use condition. However, it doesn't explicitly state when NOT to use it or name alternatives among sibling tools, though the context implies 'generate_image' might be for creating new images rather than editing existing ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_imageA
Generate a NEW image from a text prompt using Gemini. Use this ONLY when creating a completely new image, not when modifying an existing one.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text prompt describing the NEW image to create from scratch (max 10,000 chars) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the tool uses Gemini and creates new images, but lacks details on behavioral traits like rate limits, authentication needs, output format, or potential errors. The description doesn't contradict annotations (none exist), but it's insufficient for a mutation tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero waste: the first states the purpose and method, the second provides crucial usage guidance. It's front-loaded with the core function and efficiently distinguishes from siblings without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool (image generation) with no annotations and no output schema, the description is incomplete. It covers purpose and usage well but lacks behavioral context like what the output looks like, error conditions, or limitations. The schema handles the parameter, but overall completeness is moderate given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single 'prompt' parameter with its type, description, and constraints. The description adds no additional parameter semantics beyond what's in the schema, such as prompt formatting tips or examples. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Generate a NEW image'), resource ('image'), and method ('from a text prompt using Gemini'). It explicitly distinguishes this tool from its siblings by specifying 'creating a completely new image, not when modifying an existing one,' which directly contrasts with 'continue_editing' and 'edit_image' tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool ('ONLY when creating a completely new image') and when not to use it ('not when modifying an existing one'), with clear alternatives implied by the sibling tool names ('continue_editing' and 'edit_image'). This gives the agent precise context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v1.0.0- First observed
continue_editing - First observed
edit_image - First observed
generate_image
TDQS
Each tool has a clearly distinct purpose: generate_image creates new images, edit_image modifies existing images with a file path, and continue_editing iteratively edits the last image without a file path. The descriptions explicitly differentiate when to use each tool, eliminating any ambiguity.
All tool names follow a consistent verb_noun pattern (generate_image, edit_image, continue_editing), using snake_case throughout. The naming is predictable and aligns with the actions they perform, making the set easy to understand at a glance.
With 3 tools, this server is well-scoped for image generation and editing tasks. Each tool earns its place by covering distinct aspects of the workflow: creation, modification with a file, and iterative editing without a file, avoiding redundancy or gaps.
The tool set covers core image operations: generate, edit with a file, and continue editing without a file. A minor gap is the lack of tools for deleting or managing images, but this is reasonable for a focused image editing server, as agents can work around this with external file handling.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
MCP server for Google Veo AI video generation
MCP server for Grok Imagine AI video generation
MCP server for NanoBanana AI image generation and editing
Related MCP Servers
- FlicenseAqualityDmaintenanceAn MCP server that enables AI image generation, editing, and analysis using Google's Gemini 3.0 models. It supports high-resolution outputs up to 4K, style transfers, and multi-image mixing through specialized tools.3-
- AlicenseAqualityDmaintenanceAn MCP server that provides AI image generation and editing capabilities using Google's Gemini 2.5 Flash Image API. It allows users to create new images from text, modify existing files, and perform iterative edits through natural language prompts.6758MIT
- AlicenseAqualityCmaintenanceA privacy-conscious MCP server that enables image analysis via Google Gemini with safety features like two-step confirmation and SSRF defenses.2MIT
- AlicenseAqualityDmaintenanceMCP server for Google Gemini image generation with configurable model support, enabling text-to-image generation, image editing, and iterative refinement.641MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/DojoCodingLabs/nanobanana-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server