gemini-webapi-mcp
This server provides a free, no-API-key MCP interface to Google Gemini using browser cookies for authentication. Here's what you can do:
Generate images from text prompts (Nano Banana 2 model), with support for non-square aspect ratios, automatic 2x upscaling, and built-in watermark removal via Reverse Alpha Blending.
Edit existing images by providing an image file path alongside an editing instruction (e.g., "make the background blue").
Iteratively refine images by passing a
conversation_idfrom a previous generation to continue in the same session.Chat with Gemini in single or multi-turn text conversations, supporting Flash, Pro, and Flash-Thinking models.
Start persistent chat sessions with
gemini_start_chatto maintain context across multiple messages.Analyze files (images, PDFs, documents, videos) by uploading them and asking questions about their content.
Analyze URLs — including YouTube videos, webpages, and articles — for summaries or Q&A.
Reset the client (
gemini_reset) to re-initialize authentication or clear state when errors occur.Auto-authenticate via Chrome browser cookies, or manually provide
__Secure-1PSIDand__Secure-1PSIDTScookies. Supports multiple Google accounts and configurable response language.
Provides tools for generating and editing images, analyzing files (video, PDF, images, documents), and conducting text chats using Google Gemini models, with support for multi-turn conversations and auto-removal of watermarks.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-webapi-mcpgenerate an image of a cat in watercolor style"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Возможности
Генерация изображений по текстовому описанию (Nano Banana 2 с поддержкой пропорций)
2x разрешение — автоматически скачивает upscaled-версию (2048x2048 → 2816x1536 и выше)
Редактирование изображений — отправьте картинку + промпт и получите изменённую версию
Анализ файлов — видео, изображения, PDF, документы
Текстовый чат с Gemini (Flash, Pro, Flash-Thinking)
Авто-удаление вотермарки — sparkle-метка Gemini убирается встроенным детерминированным Reverse Alpha Blending (без внешних бинарников, ML-моделей и скачиваний)
Авто-аутентификация через cookies из Chrome
Related MCP server: Nano Banana MCP
Быстрый старт
1. Войдите в Gemini
Откройте Chrome, перейдите на gemini.google.com и войдите в свой Google-аккаунт.
2. Установите MCP-сервер
Из GitHub (без клонирования):
uv run --with "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git" gemini-webapi-mcpЛокальная установка:
git clone https://github.com/AndyShaman/gemini-webapi-mcp.git
cd gemini-webapi-mcp
uv sync
uv run gemini-webapi-mcp3. Добавьте MCP-конфиг
claude mcp add-json gemini '{"command":"uv","args":["run","--with","gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git","gemini-webapi-mcp"]}'Или добавьте вручную в .mcp.json в корне проекта:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Добавьте в конфиг Claude Desktop:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Используйте стандартный MCP stdio-конфиг:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Путь к файлу конфига зависит от вашего MCP-клиента.
"args": ["--directory", "/path/to/gemini-webapi-mcp", "run", "gemini-webapi-mcp"]4. Установите скилл для Claude Code (опционально)
Папка skill/ содержит скилл для Claude Code — подсказки по промптингу, документацию по тулам и гайд по генерации изображений. Скилл автоматически активируется при работе с Gemini.
cp -r skill ~/.claude/skills/gemini-mcp5. Проверьте
Запустите сервер вручную — если инициализация прошла без ошибок, всё работает:
uv run --with "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git" gemini-webapi-mcpПосле этого откройте Claude Code или Claude Desktop и попробуйте: «Сгенерируй картинку кота в акварельном стиле через Gemini».
Аутентификация
Сервер автоматически читает cookies из Chrome через browser-cookie3.
Несколько Google-аккаунтов? Установите
GEMINI_ACCOUNT_INDEX— номер аккаунта из Chrome (0 = первый, 1 = второй, ...). Посмотрите порядок: кликните на аватарку в gemini.google.com.
Если автоопределение cookies не работает, задайте их вручную:
Откройте Chrome DevTools на gemini.google.com → Application → Cookies
Скопируйте значения
__Secure-1PSIDи__Secure-1PSIDTSДобавьте в MCP-конфиг:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"],
"env": {
"GEMINI_PSID": "your__Secure-1PSID_value",
"GEMINI_PSIDTS": "your__Secure-1PSIDTS_value"
}
}
}
}Переменные окружения
Переменная | Описание | По умолчанию |
| Значение cookie | авто из Chrome |
| Значение cookie | авто из Chrome |
| Язык ответов Gemini ( |
|
| Индекс Google-аккаунта (0, 1, 2, ...) |
|
Высокое разрешение (2x)
Сервер автоматически запрашивает у Google увеличенную версию сгенерированного изображения — тот же механизм, что использует кнопка "Download" в веб-интерфейсе Gemini. Google выполняет server-side upscale, и вы получаете изображение в 2x разрешении:
Модель | Нативное | 2x (скачивается) |
Flash-Thinking (16:9) | 1408x768 | 2816x1536 |
Flash-Thinking (9:16) | 768x1376 | 1536x2752 |
Flash-Thinking (1:1) | 1024x1024 | 2048x2048 |
Если 2x-версия недоступна (таймаут, ошибка сети), сервер автоматически использует нативное разрешение.
Удаление вотермарки
Gemini добавляет sparkle-метку (четырёхконечную звёздочку) в правый нижний угол сгенерированных изображений. Сервер убирает её встроенным reverse-alpha-вычитанием — original = (L − A·shape·255) / (1 − A·shape) — без внешних бинарников, ML-моделей и скачиваний.
Метка всегда стоит в одном из двух фиксированных угловых якорей — на отступе 96 px или 32 px от правого нижнего угла (абсолютные отступы, не зависят от разрешения; логотип 48 px, ×2 при 2x-upscale). Сервер не детектирует метку по порогу корреляции — это ненадёжно на ярком фоне, где слабая полупрозрачная звёздочка почти не контрастирует. Вместо этого он детерминированно обрабатывает оба якоря:
per-image прозрачность — сила метки оценивается из самого кадра (
L = bg + A·shape·(255−bg), least-squares); на пустом якореA ≈ 0, поэтому вычитание становится no-op;reverse-alpha в позиции якоря;
self-check по |corr| — вычитание принимается, только если оно уменьшает корреляцию с формой звезды; иначе откатывается. Навредить пустому, текстурному или цветному-контентному якорю невозможно.
Подход не содержит ни одной захардкоженной «под разрешение» величины, поэтому работает на любом нативном разрешении и соотношении сторон — на светлом, тёмном и цветном фоне.
Единственный калиброванный инвариант — форма звезды в src/gemini_webapi_mcp/assets/wm_alpha_edit.npy. Чтобы временно отключить удаление (получить «сырой» водяной знак), задайте GEMINI_WM_KEEP=1.
Инструменты
Инструмент | Описание |
| Генерация новых или редактирование существующих изображений |
| Анализ файлов — видео, изображения, PDF, документы |
| Анализ URL — YouTube-видео, веб-страницы, статьи |
| Текстовый чат (одиночный или multi-turn) |
| Начать multi-turn сессию |
| Переинициализация клиента при ошибках авторизации |
Модели
Модель | По умолчанию для | Примечание |
| чат, анализ файлов | Быстрая |
| генерация изображений | Nano Banana 2, поддержка пропорций |
| — | Альтернативная модель |
Примеры использования
После настройки MCP-конфига Claude сам вызывает нужные инструменты. Просто попросите в чате:
Задача | Что написать Claude |
Сгенерировать изображение | «Сгенерируй через Gemini кота в акварельном стиле» |
Отредактировать изображение | «Отредактируй через Gemini /path/to/cat.png — сделай кота серым» |
Итеративная правка | «Теперь сделай фон темнее» (в том же разговоре) |
Проанализировать видео | «Проанализируй через Gemini это видео: https://youtube.com/watch?v=...» |
Проанализировать файл | «Загрузи в Gemini /path/to/doc.pdf и сделай краткое резюме» |
Инструменты, которые Claude вызовет:
gemini_generate_image(prompt="кот в акварельном стиле")
gemini_generate_image(prompt="сделай кота серым", files=["/path/to/cat.png"])
gemini_generate_image(prompt="сделай фон темнее", conversation_id=["c_abc", "r_123", "rc_456"])
gemini_analyze_url(url="https://youtube.com/watch?v=...", prompt="О чём это видео?")
gemini_upload_file(file_path="/path/to/doc.pdf", prompt="Сделай краткое резюме")Благодарности
Этот проект построен на основе библиотеки gemini-webapi от @HanaokaYuzu (форк @xob0t с поддержкой curl_cffi) — реверс-инжиниринговой асинхронной Python-обёртки для веб-приложения Google Gemini. Лицензия: AGPL-3.0.
Алгоритм удаления вотермарки (Reverse Alpha Blending) изначально вдохновлён проектами gwt-mini от @allenk (Allen Kuo, MIT License) и gemini-watermark-remover от @GargantuaX (MIT License). В текущей версии сервер использует собственную встроенную реализацию и откалиброванные alpha-карты — внешние бинарники не требуются.
Лицензия
AGPL-3.0 — свободно используйте, модифицируйте и распространяйте при условии сохранения открытости исходного кода.
@AndyShaman · gemini-webapi-mcp
Features
Image generation from text descriptions (Nano Banana 2 with aspect ratio support)
2x resolution — automatically downloads upscaled version (2048x2048 → 2816x1536 and above)
Image editing — send an image + prompt to get a modified version
File analysis — video, images, PDF, documents
Text chat with Gemini (Flash, Pro, Flash-Thinking)
Auto watermark removal — Gemini's sparkle mark is stripped by a built-in deterministic Reverse Alpha Blending pass (no external binaries, ML models, or downloads)
Auto-authentication via Chrome browser cookies
Quick Start
1. Log into Gemini
Open Chrome, go to gemini.google.com and sign in.
2. Install the MCP server
From GitHub (no clone needed):
uv run --with "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git" gemini-webapi-mcpLocal install:
git clone https://github.com/AndyShaman/gemini-webapi-mcp.git
cd gemini-webapi-mcp
uv sync
uv run gemini-webapi-mcp3. Add MCP config
claude mcp add-json gemini '{"command":"uv","args":["run","--with","gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git","gemini-webapi-mcp"]}'Or add manually to .mcp.json in your project root:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Add to Claude Desktop config:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Use the standard MCP stdio config:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"]
}
}
}Config file path depends on your MCP client.
"args": ["--directory", "/path/to/gemini-webapi-mcp", "run", "gemini-webapi-mcp"]4. Install the skill for Claude Code (optional)
The skill/ folder contains a Claude Code skill — prompting tips, tool documentation and an image generation guide. The skill auto-activates when working with Gemini.
cp -r skill ~/.claude/skills/gemini-mcp5. Verify
Run the server manually — if it initializes without errors, everything works:
uv run --with "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git" gemini-webapi-mcpThen open Claude Code or Claude Desktop and try: "Generate a watercolor cat image with Gemini".
Authentication
The server reads cookies from Chrome automatically via browser-cookie3.
Multiple Google accounts? Set
GEMINI_ACCOUNT_INDEX— the account number from Chrome (0 = first, 1 = second, ...). Check the order by clicking your avatar on gemini.google.com.
If cookie auto-detection fails, set them manually:
Open Chrome DevTools on gemini.google.com → Application → Cookies
Copy
__Secure-1PSIDand__Secure-1PSIDTSvaluesAdd to your MCP config:
{
"mcpServers": {
"gemini": {
"command": "uv",
"args": ["run", "--with", "gemini-webapi-mcp @ git+https://github.com/AndyShaman/gemini-webapi-mcp.git", "gemini-webapi-mcp"],
"env": {
"GEMINI_PSID": "your__Secure-1PSID_value",
"GEMINI_PSIDTS": "your__Secure-1PSIDTS_value"
}
}
}
}Environment Variables
Variable | Description | Default |
| Cookie value | auto from Chrome |
| Cookie value | auto from Chrome |
| Gemini response language ( |
|
| Google account index (0, 1, 2, ...) |
|
High Resolution (2x)
The server automatically requests an upscaled version of each generated image — the same mechanism used by the "Download" button in Gemini's web interface. Google performs server-side upscaling, delivering images at 2x resolution:
Model | Native | 2x (downloaded) |
Flash-Thinking (16:9) | 1408x768 | 2816x1536 |
Flash-Thinking (9:16) | 768x1376 | 1536x2752 |
Flash-Thinking (1:1) | 1024x1024 | 2048x2048 |
If the 2x version is unavailable (timeout, network error), the server automatically falls back to native resolution.
Watermark Removal
Gemini adds a sparkle watermark (4-point star) to the bottom-right corner of generated images. The server removes it with a built-in reverse alpha-blend — original = (L − A·shape·255) / (1 − A·shape) — with no external binaries, ML models, or downloads.
The mark is always stamped at one of two fixed corner anchors — 96 px or 32 px from the bottom-right (absolute offsets that don't depend on resolution; the logo is 48 px, ×2 when 2x-upscaled). The server does not detect the mark by a correlation threshold — that's unreliable on bright backgrounds, where a faint translucent star barely contrasts. Instead it deterministically processes both anchors:
per-image opacity — the mark's strength is fit from the frame itself (
L = bg + A·shape·(255−bg), least-squares); on an empty anchorA ≈ 0, so the subtraction is a no-op;reverse alpha-blend at the anchor;
|corr| self-check — a subtraction is accepted only if it reduces the correlation with the star shape; otherwise it is reverted. It can never damage an empty, textured, or coloured-content anchor.
The approach hard-codes no resolution-dependent values, so it works at any native resolution or aspect ratio — on light, dark, and coloured backgrounds alike.
The only calibrated invariant is the star shape in src/gemini_webapi_mcp/assets/wm_alpha_edit.npy. Set GEMINI_WM_KEEP=1 to disable removal (keep the raw watermark).
Tools
Tool | Description |
| Generate new or edit existing images |
| Analyze files — video, images, PDF, documents |
| Analyze URLs — YouTube videos, webpages, articles |
| Text chat (single or multi-turn) |
| Start a multi-turn session |
| Re-initialize client on auth errors |
Models
Model | Default for | Notes |
| chat, file analysis | Fast |
| image generation | Nano Banana 2, supports aspect ratios |
| — | Alternative model |
Usage Examples
Once configured, Claude calls the right tools automatically. Just ask in chat:
Task | What to tell Claude |
Generate an image | "Generate a watercolor cat with Gemini" |
Edit an image | "Edit /path/to/cat.png with Gemini — make the cat gray" |
Iterative refinement | "Now make the background darker" (same conversation) |
Analyze a video | "Analyze this video with Gemini: https://youtube.com/watch?v=..." |
Analyze a file | "Upload /path/to/doc.pdf to Gemini and summarize it" |
Tools that Claude will call:
gemini_generate_image(prompt="a cat in watercolor style")
gemini_generate_image(prompt="make it gray", files=["/path/to/cat.png"])
gemini_generate_image(prompt="make the background darker", conversation_id=["c_abc", "r_123", "rc_456"])
gemini_analyze_url(url="https://youtube.com/watch?v=...", prompt="Summarize this video")
gemini_upload_file(file_path="/path/to/doc.pdf", prompt="Summarize key points")Acknowledgements
This project is built on top of gemini-webapi by @HanaokaYuzu (fork by @xob0t with curl_cffi support) — a reverse-engineered async Python wrapper for the Google Gemini web app. Licensed under AGPL-3.0.
The watermark-removal algorithm (Reverse Alpha Blending) was originally inspired by gwt-mini by @allenk (Allen Kuo, MIT License) and gemini-watermark-remover by @GargantuaX (MIT License). The current version uses its own built-in implementation and calibrated alpha maps — no external binaries required.
License
AGPL-3.0 — free to use, modify, and distribute, provided the source code remains open.
Available Tools
6 toolsgemini_analyze_urlARead-only
Analyze a URL — YouTube videos, webpages, articles, etc.
Gemini can watch YouTube videos and read webpages, then answer questions about their content.
Args: url: The URL to analyze (YouTube, article, webpage, etc.). prompt: Question or instruction about the content (e.g. 'Summarize this video', 'What are the key points?'). model: Model name. Defaults to gemini-3.0-flash.
Returns: Gemini's analysis of the URL content.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| prompt | No | Summarize this content. | |
| model | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond annotations by stating 'Gemini can watch YouTube videos and read webpages, then answer questions.' This clarifies the tool's capabilities (including media processing). Annotations (readOnlyHint: true, destructiveHint: false) are consistent and the description adds value without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a two-sentence overview followed by a structured Args section. Every sentence adds value, and the most critical information (purpose and capability) is front-loaded. No unnecessary text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity and presence of an output schema, the description covers the tool's operation well (Args, returns, examples of content types). It lacks details on auth or rate limits, but annotations handle basic safety. Nearly complete for standard usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description includes an Args section that explains each parameter's purpose (e.g., 'url: The URL to analyze', 'prompt: Question or instruction'), adding substantial meaning beyond the input schema's type/default information. With 0% schema description coverage, this fully compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Analyze a URL' and elaborates with specific examples like YouTube videos and webpages. It uses a specific verb ('analyze') and resource ('URL'), effectively distinguishing it from sibling tools like gemini_chat or gemini_generate_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (analyzing URL content) and implies it should be used instead of other tools for content analysis. However, it does not explicitly state when not to use it or mention alternatives, missing some guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_chatARead-only
Send a text prompt to Google Gemini and get a response.
Args: prompt: The text prompt to send to Gemini. model: Model name (e.g. 'gemini-3.0-flash', 'gemini-3.0-pro', 'gemini-3.0-flash-thinking'). Defaults to gemini-3.0-flash. session_id: Optional session ID from gemini_start_chat for multi-turn conversation with context.
Returns: Gemini's text response. When using flash-thinking model, also includes the model's reasoning process.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| model | No | ||
| session_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond annotations: it states that the output includes the model's reasoning process when using flash-thinking model. Annotations already indicate readOnlyHint=true and destructiveHint=false, so no contradiction. The description is transparent about the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a single paragraph with 'Args' and 'Returns' sections. Every sentence adds value, and the main action is front-loaded. No unnecessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (3 parameters, output schema exists, annotations present), the description is complete. It covers parameter semantics, return value, and special behavior (thinking model reasoning), providing all necessary context for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully explains all three parameters: prompt (text to send), model (with examples and default), and session_id (for multi-turn context from gemini_start_chat). This adds significant meaning beyond the schema's basic titles and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Send a text prompt to Google Gemini and get a response.' It specifies the verb (send), resource (text prompt to Gemini), and outcome (get response). It also distinguishes from sibling tools like gemini_start_chat by mentioning session_id for multi-turn conversation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage guidance, such as using the optional session_id from gemini_start_chat for multi-turn conversation and listing model options. However, it could be more explicit about when to use this tool compared to alternatives like gemini_analyze_url. The guidelines are mostly clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_generate_imageA
Generate or edit images with Gemini.
Without files: generates a new image from the text prompt. With files: edits/transforms the provided image(s) based on the prompt.
Pass conversation_id from a previous call to continue refining images in the same conversation thread (e.g. "make it more dramatic", "add rain"). You can also use a cid from the Gemini web URL (gemini.google.com/app/{cid}).
Images are saved to ~/Pictures/gemini/ and full file paths are returned.
Args: prompt: Description of the image to generate, or editing instruction (e.g. 'change the background to blue', 'make it a cartoon'). model: Model name. Defaults to gemini-3.0-flash-thinking (Nano Banana 2, supports non-square aspect ratios). files: Optional list of file paths to images to edit/transform. conversation_id: Optional list of [cid, rid, rcid] from a previous gemini_generate_image response to continue the conversation. Passing just [cid] (from browser URL) also works.
Returns: JSON with generated image paths, conversation_id for continuation, or an error message.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| model | No | ||
| files | No | ||
| conversation_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that images are saved to ~/Pictures/gemini/ and full paths are returned. Explains conversation continuation behavior. Annotations (readOnlyHint=false, destructiveHint=false) are not contradicted; description adds meaningful context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with clear sections for modes, args, and returns. Every sentence adds value; no redundancy. Concise yet comprehensive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (image generation, editing, conversation history) and the presence of an output schema, the description covers all necessary aspects: mode selection, parameter usage, file saving, and return format. An agent can effectively invoke the tool based on this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description provides detailed explanations for all parameters: prompt (generation vs editing), model (name, default, aspect ratio hint), files (optional paths), and conversation_id (continuation method). Adds significant meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Generate or edit images with Gemini' and distinguishes between two modes (without files for new images, with files for editing). It uniquely identifies the tool's purpose among siblings, which are chat/analyze/upload/reset tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly guides when to use without files (new image) and with files (edit/transform). Also provides instructions for using conversation_id to continue refinement, and mentions model defaults. No direct alternatives exist among siblings, so this is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_resetAIdempotent
Re-initialise the Gemini client (refresh cookies, clear state).
Use this when you get authentication errors or want a fresh session.
Returns: Confirmation message or error.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses what the tool does (refresh cookies, clear state) beyond annotations. Annotations show idempotentHint=true and destructiveHint=false, which align with description. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: purpose, usage scenario, return value. No wasted words. Front-loaded with key action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and a simple action, the description fully explains purpose, when to use, and what to expect. Output schema is not needed as return type is described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist; schema coverage is 100%. Description does not need to add param info. Baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action: 'Re-initialise the Gemini client' with specific actions (refresh cookies, clear state). It is distinct from sibling tools like gemini_chat or gemini_analyze_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises when to use: 'when you get authentication errors or want a fresh session'. Does not mention when not to use alternatives, but for a single-purpose tool this is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_start_chatA
Start a new multi-turn chat session with Gemini.
The session maintains conversation history so follow-up messages have full context. Pass the returned session_id to gemini_chat.
Args: model: Model name for this session. Defaults to gemini-3.0-flash.
Returns: JSON with session_id to use in subsequent gemini_chat calls.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations set readOnlyHint false, destructiveHint false, etc. Description adds that session maintains history and returns session_id. Discloses behavioral trait of context preservation. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise – three sentences plus structured Args/Returns sections. Front-loaded with main purpose. No redundant words. Each sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 optional param, returns session_id), the description fully covers usage: what it does, how to use the output, and the model parameter. Sufficient for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter 'model' has 0% schema description coverage. Description provides default value (gemini-3.0-flash) and explanation, adding meaning beyond the schema. Could list available models, but adequate for a default parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Start a new multi-turn chat session with Gemini' with specific verb and resource. Distinguishes itself from siblings like gemini_chat and gemini_reset by indicating it initiates a session that maintains history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains that the session maintains conversation history for context, and instructs to pass the returned session_id to gemini_chat. Implicitly guides when to use this tool. Could explicitly state when not to use, but sufficient for typical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_upload_fileARead-only
Upload a file (image, PDF, document, video) to Gemini and ask a question about it.
Args: file_path: Absolute path to the file to upload. prompt: Question or instruction about the file (e.g. 'What is shown in this image?'). model: Model name. Defaults to gemini-3.0-flash.
Returns: Gemini's text response about the uploaded file.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| prompt | No | Describe this file. | |
| model | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description claims to upload a file, which is a write operation, but the annotations set readOnlyHint=true, creating a direct contradiction. This misleads the agent about the tool's side effects. Beyond this, the description only adds basic return info and default model, but fails to disclose rate limits, file size constraints, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with separate args and returns sections. It is front-loaded with the purpose and remains concise, though the args section could be slightly streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and the tool's moderate complexity, the description covers the basic workflow. However, it omits important contextual details like supported file formats, size limits, and error handling, which are critical for robust agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description provides essential meaning for all three parameters, including the default for model. However, the model parameter's default in the schema is null, while the description states a specific model name ('gemini-3.0-flash'), creating a discrepancy that could confuse agents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'upload' with the resource 'file to Gemini' and specifies the action of asking a question. This distinguishes it from sibling tools like gemini_chat or gemini_analyze_url, which handle different modalities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly suggests using the tool when a file needs analysis, but it does not explicitly contrast with siblings or provide when-not conditions. The purpose is clear enough from context, but explicit usage guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
6 tool updates
v0.1.0- First observed
gemini_analyze_url - First observed
gemini_chat - First observed
gemini_generate_image - First observed
gemini_reset - First observed
gemini_start_chat - First observed
gemini_upload_file
TDQS
Each tool has a unique function: analyzing URLs, chatting, generating images, resetting the client, starting multi-turn sessions, and uploading files. There is no overlap or ambiguity in their purposes.
All tools use a consistent 'gemini_' prefix and snake_case, but there is slight inconsistency in verb vs. verb_noun patterns (e.g., 'gemini_chat' vs 'gemini_start_chat'). The naming is mostly predictable and descriptive.
With 6 tools, the set is well-scoped without being too few or too many. Each tool serves a distinct purpose, covering core interactions with Gemini.
The tool set covers all major interaction modes: text chat, image generation, URL analysis, file upload, session management, and client reset. Minor gaps like model listing or conversation history management are non-essential.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Google Veo AI video generation
MCP server for Qwen Image 3 AI image generation
MCP server for NanoBanana AI image generation and editing
MCP server for Hailuo (MiniMax) AI video generation
Related MCP Servers
- AlicenseAqualityBmaintenanceMCP server for Google Gemini image generation, editing, and processing, with two tools and no bloat.22671MIT
- AlicenseAqualityDmaintenanceMCP server for Google Gemini image generation with configurable model support, enabling text-to-image generation, image editing, and iterative refinement.641MIT
- AlicenseAqualityDmaintenanceMCP server for generating and editing images using Google Gemini API, with support for multi-turn iterative refinement.322MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for AI image generation and editing using Google Gemini image models.758MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/AndyShaman/gemini-webapi-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server