Skip to main content
Glama

nanobanana-mcp

A hardened MCP server for Gemini image generation. Fork of ConechoAI/Nano-Banana-MCP with security fixes, strict TypeScript, and model selection.

Features

  • 3 tools: generate_image, edit_image, continue_editing

  • Model selection via NANOBANANA_MODEL env var with whitelist validation

  • Security hardened: path traversal protection, file size limits, no plaintext key storage

  • Strict TypeScript: zero any types, Zod validation on all inputs

Related MCP server: Nano-Banana MCP Server

Quick Start

Claude Code

Add to ~/.claude/settings.json:

{
  "mcpServers": {
    "nanobanana": {
      "command": "npx",
      "args": ["tsx", "/path/to/nanobanana-mcp/src/index.ts"],
      "env": {
        "GEMINI_API_KEY": "your-api-key",
        "NANOBANANA_MODEL": "gemini-2.5-flash-image"
      }
    }
  }
}

Other MCP Clients

GEMINI_API_KEY=your-key npx tsx src/index.ts

The server communicates over stdio using the MCP protocol.

Tools

generate_image

Generate a new image from a text prompt.

prompt (required): Text describing the image to create (max 10,000 chars)

edit_image

Edit an existing image with a text prompt.

imagePath (required): Full file path to the image to edit
prompt (required): Text describing the modifications (max 10,000 chars)
referenceImages (optional): Array of file paths to reference images

continue_editing

Continue editing the last generated/edited image in the current session.

prompt (required): Text describing changes to make (max 10,000 chars)
referenceImages (optional): Array of file paths to reference images

Configuration

All configuration is via environment variables. No config files are written to disk.

Variable

Required

Description

GEMINI_API_KEY

Yes

Google Gemini API key

NANOBANANA_GEMINI_API_KEY

No

Override for GEMINI_API_KEY (takes priority)

NANOBANANA_MODEL

No

Model to use (see below)

Available Models

Model ID

Description

gemini-2.5-flash-image

Fast generation, good for high-volume use (default)

gemini-3-pro-image-preview

Pro quality, complex prompts, better text rendering

gemini-3.1-flash-image-preview

Latest model, advanced features

Output

Generated images are saved to ~/nanobanana-images/ with unique filenames. The tool response includes both the file path and the image data inline.

Security

This fork addresses the following security issues from the original:

Issue

Fix

API key saved to disk in plaintext

Removed config file persistence entirely

configure_gemini_token tool accepts key via MCP

Tool removed; keys only via env vars

Path traversal in editImage

validatePath() checks paths resolve within $HOME or $TMPDIR

No prompt length validation

Capped at 10,000 chars via Zod

Hardcoded model

NANOBANANA_MODEL env var with whitelist

Silent swallowing of reference image errors

Errors now thrown and reported

Math.random() for filenames

crypto.randomUUID()

No file size limit on reads

Max 20MB

Verbose errors leak internal paths

Sanitized error messages

process.cwd() fallback for output dir

Fixed to ~/nanobanana-images/

Development

npm install
npm run typecheck   # Type check without emitting
npm run dev         # Run with tsx (hot reload)
npm run build       # Compile to dist/

Project Structure

src/
  index.ts          # MCP server entry point (3 tool handlers)
  gemini-client.ts  # Gemini API wrapper with model selection
  file-handler.ts   # Secure file I/O with path validation
  types.ts          # TypeScript interfaces and Zod schemas

License

MIT - Based on ConechoAI/Nano-Banana-MCP

Available Tools

3 tools
continue_editingA

Continue editing the LAST image generated or edited in this session. Automatically uses the previous image without needing a file path. Use for iterative improvements.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText describing changes to make to the last image (max 10,000 chars)
referenceImagesNoOptional array of file paths to reference images

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses key behavioral traits: it operates on session state ('LAST image... in this session'), is stateful ('Automatically uses the previous image'), and is designed for 'iterative improvements'. However, it doesn't mention potential limitations like session duration, error handling, or what happens if no previous image exists, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by operational details and usage context. Every sentence earns its place: the first defines the tool, the second explains automation, and the third provides guidance. It's concise with zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (stateful operation, 2 parameters), no annotations, and no output schema, the description is largely complete. It covers purpose, usage, and key behavior. However, it lacks details on output format or error cases, which could be helpful for an agent. Since there's no output schema, some additional context on returns would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds no specific parameter semantics beyond implying 'prompt' is for 'changes to make' and 'referenceImages' might be optional references. This meets the baseline of 3 since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Continue editing'), the resource ('the LAST image generated or edited in this session'), and distinguishes it from siblings by emphasizing it works on the 'previous image without needing a file path' unlike edit_image which likely requires a file path. The phrase 'iterative improvements' further clarifies its unique role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Continue editing the LAST image... in this session') and provides clear alternatives by naming sibling tools (edit_image, generate_image) in the context. It specifies 'Automatically uses the previous image' which implies when not to use it (e.g., when starting fresh or editing a different image).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_imageB

Edit an existing image file with a text prompt, optionally using additional reference images. Use this when you have the exact file path of an image to modify.

ParametersJSON Schema
NameRequiredDescriptionDefault
imagePathYesFull file path to the image to edit
promptYesText describing the modifications to make (max 10,000 chars)
referenceImagesNoOptional array of file paths to reference images (for style transfer, adding elements, etc.)

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While it mentions the tool edits images with a prompt and optional references, it doesn't disclose critical behavioral traits like whether this is a destructive operation (overwrites the original file?), what permissions are needed, rate limits, output format, or error conditions. For a mutation tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and well-structured: two sentences that efficiently convey purpose and usage guidelines. Every word earns its place with no redundancy or fluff. It's front-loaded with the core functionality followed by the key usage condition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with 3 parameters, no annotations, and no output schema, the description is incomplete. It adequately covers purpose and basic usage but lacks crucial behavioral context (destructive nature, permissions, output format) and doesn't compensate for the absence of structured metadata. The agent would be left guessing about important operational aspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds minimal value beyond what's in the schema: it mentions 'text prompt' and 'reference images' but doesn't provide additional semantic context like examples, constraints beyond the schema's max length, or how reference images are used. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Edit an existing image file with a text prompt, optionally using additional reference images.' It specifies the verb ('edit'), resource ('existing image file'), and key mechanisms (text prompt, reference images). However, it doesn't explicitly distinguish this from sibling tools like 'continue_editing' or 'generate_image' beyond mentioning file path requirements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance: 'Use this when you have the exact file path of an image to modify.' This gives a specific when-to-use condition. However, it doesn't explicitly state when NOT to use it or name alternatives among sibling tools, though the context implies 'generate_image' might be for creating new images rather than editing existing ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate a NEW image from a text prompt using Gemini. Use this ONLY when creating a completely new image, not when modifying an existing one.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText prompt describing the NEW image to create from scratch (max 10,000 chars)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the tool uses Gemini and creates new images, but lacks details on behavioral traits like rate limits, authentication needs, output format, or potential errors. The description doesn't contradict annotations (none exist), but it's insufficient for a mutation tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero waste: the first states the purpose and method, the second provides crucial usage guidance. It's front-loaded with the core function and efficiently distinguishes from siblings without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool (image generation) with no annotations and no output schema, the description is incomplete. It covers purpose and usage well but lacks behavioral context like what the output looks like, error conditions, or limitations. The schema handles the parameter, but overall completeness is moderate given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single 'prompt' parameter with its type, description, and constraints. The description adds no additional parameter semantics beyond what's in the schema, such as prompt formatting tips or examples. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Generate a NEW image'), resource ('image'), and method ('from a text prompt using Gemini'). It explicitly distinguishes this tool from its siblings by specifying 'creating a completely new image, not when modifying an existing one,' which directly contrasts with 'continue_editing' and 'edit_image' tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool ('ONLY when creating a completely new image') and when not to use it ('not when modifying an existing one'), with clear alternatives implied by the sibling tool names ('continue_editing' and 'edit_image'). This gives the agent precise context for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv1.0.0
    • First observedcontinue_editing
    • First observededit_image
    • First observedgenerate_image

TDQS

A4.1/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: generate_image creates new images, edit_image modifies existing images with a file path, and continue_editing iteratively edits the last image without a file path. The descriptions explicitly differentiate when to use each tool, eliminating any ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (generate_image, edit_image, continue_editing), using snake_case throughout. The naming is predictable and aligns with the actions they perform, making the set easy to understand at a glance.

Tool Count5/5

With 3 tools, this server is well-scoped for image generation and editing tasks. Each tool earns its place by covering distinct aspects of the workflow: creation, modification with a file, and iterative editing without a file, avoiding redundancy or gaps.

Completeness4/5

The tool set covers core image operations: generate, edit with a file, and continue editing without a file. A minor gap is the lack of tools for deleting or managing images, but this is reasonable for a focused image editing server, as agents can work around this with external file handling.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/DojoCodingLabs/nanobanana-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server