Skip to main content
Glama

MCP Fooocus API

A Model Context Protocol (MCP) server that provides text-to-image generation capabilities through the Fooocus Stable Diffusion API.

Features

  • Text-to-Image Generation: Generate high-quality images from text prompts

  • Intelligent Style Selection: Automatically selects 1-3 appropriate styles based on your prompt

  • Custom Style Override: Manually specify styles from 300+ available options

  • Multiple Performance Modes: Choose between Speed, Quality, and Extreme Speed

  • Configurable Aspect Ratios: Support for various image dimensions

  • Environment-based Configuration: Easy API endpoint configuration via .env file

Related MCP server: Volcengine Image Generation MCP Server

Installation

Install directly from GitHub:

uv add git+https://github.com/raihan0824/mcp-fooocus-api.git

Or install from PyPI (when published):

uv add mcp-fooocus-api

Run with uvx:

uvx --from git+https://github.com/raihan0824/mcp-fooocus-api.git mcp-fooocus-api

Using pip

pip install git+https://github.com/raihan0824/mcp-fooocus-api.git

Or from PyPI (when published):

pip install mcp-fooocus-api

Development Installation

# Clone the repository
git clone https://github.com/raihan0824/mcp-fooocus-api.git
cd mcp-fooocus-api

# Install with uv
uv sync --dev

# Or install with pip
pip install -e ".[dev]"

Configuration

  1. Copy the example environment file:

cp .env.example .env
  1. Edit the .env file to configure your Fooocus API endpoint:

FOOOCUS_API_URL=http://103.125.100.56:8888/v1/generation/text-to-image

Usage

Available Tools

The MCP server provides three main tools:

1. generate_image

Generate an image using the Fooocus API.

Parameters:

  • prompt (required): Text description of the image to generate

  • performance (optional): Performance setting - "Speed" (default), "Quality", or "Extreme Speed"

  • custom_styles (optional): Comma-separated list of custom styles

  • aspect_ratio (optional): Image dimensions (default: "1024*1024")

Example:

{
  "prompt": "A serene landscape with mountains and a lake at sunset",
  "performance": "Quality",
  "aspect_ratio": "1024*1024"
}

2. list_available_styles

Lists all available styles organized by category.

Returns:

  • Total number of available styles

  • Styles organized by categories (Fooocus, SAI, MRE, Art Styles, etc.)

  • Available performance options

3. get_server_info

Get information about the server configuration and capabilities.

Returns:

  • Server version and name

  • Configured API endpoint

  • Available features

  • Performance options

Style Categories

The server includes 300+ styles organized into categories:

  • Fooocus Styles: Native Fooocus styles (V2, Enhance, Sharp, etc.)

  • SAI Styles: Stability AI styles (Photographic, Digital Art, Anime, etc.)

  • Art Styles: Classical art movements (Renaissance, Impressionist, Cubist, etc.)

  • Photography: Various photography styles (Film Noir, HDR, Macro, etc.)

  • Game Styles: Video game-inspired styles (Minecraft, Pokemon, Retro, etc.)

  • Futuristic: Sci-fi and cyberpunk styles

  • And many more...

Intelligent Style Selection

When you don't specify custom styles, the server automatically selects appropriate styles based on your prompt:

  • "renaissance portrait" → Selects "Artstyle Renaissance"

  • "cyberpunk city" → Selects "Futuristic Cyberpunk Cityscape"

  • "anime character" → Selects "SAI Anime"

  • "realistic photo" → Selects "SAI Photographic"

  • "watercolor painting" → Selects "Artstyle Watercolor"

Running the Server

As an MCP Server

Add to your MCP client configuration:

{
  "mcpServers": {
    "fooocus": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/raihan0824/mcp-fooocus-api.git", "mcp-fooocus-api"]
    }
  }
}

Or if installed from PyPI:

{
  "mcpServers": {
    "fooocus": {
      "command": "uvx",
      "args": ["mcp-fooocus-api"]
    }
  }
}

Standalone Server

You can also run the server directly:

# With uv
uvx --from git+https://github.com/raihan0824/mcp-fooocus-api.git mcp-fooocus-api --port 3000 --host localhost

# Or if installed locally
python -m mcp_fooocus_api.server --port 3000 --host localhost

API Response Format

Successful generation returns:

{
  "success": true,
  "prompt": "Your prompt here",
  "selected_styles": ["Style1", "Style2"],
  "performance": "Speed",
  "aspect_ratio": "1024*1024",
  "result": {
    // Fooocus API response data
  }
}

Error responses include:

{
  "success": false,
  "error": "Error description",
  "prompt": "Your prompt here",
  "selected_styles": ["Style1", "Style2"]
}

Requirements

  • Python 3.8+

  • Access to a Fooocus API endpoint

  • Internet connection for API requests

Dependencies

  • mcp >= 1.0.0

  • httpx >= 0.27

  • python-dotenv >= 1.0.0

  • pydantic >= 2.7.2, < 3.0.0

Development

To set up for development:

  1. Clone the repository

  2. Install dependencies: pip install -e .

  3. Configure your .env file

  4. Run the server: python -m mcp_fooocus_api.server

License

MIT License - see LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Support

For issues and questions, please visit the GitHub repository.

Available Tools

3 tools
generate_imageC
Generate an image using Fooocus Stable Diffusion API.

Args:
    prompt: Text description of the image to generate
    performance: Performance setting (Speed, Quality, Extreme Speed)
    custom_styles: Comma-separated list of custom styles (optional, will auto-select if not provided)
    aspect_ratio: Image aspect ratio (default: 1024*1024)

Returns:
    Dictionary containing generation result with image URLs
ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
performanceNoSpeed
custom_stylesNo
aspect_ratioNo1024*1024

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions the API and return format but lacks critical behavioral details: no information on rate limits, authentication needs, error handling, or what happens if generation fails. For a generative tool with zero annotation coverage, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections for Args and Returns, making it easy to scan. It's concise with no wasted sentences, though it could be more front-loaded by emphasizing the core purpose first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (image generation with 4 parameters) and no annotations, the description is moderately complete. It covers parameters and return format, but lacks behavioral context and usage guidelines. The output schema exists, so describing return values isn't needed, but other gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It lists all 4 parameters with brief explanations (e.g., 'Text description of the image to generate' for prompt), adding meaning beyond the schema's titles. However, it doesn't fully detail constraints like valid aspect ratio formats or performance options, leaving gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate an image using Fooocus Stable Diffusion API.' It specifies the verb ('Generate') and resource ('image'), though it doesn't explicitly differentiate from sibling tools like list_available_styles. The purpose is specific but lacks sibling comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description mentions the API but doesn't specify scenarios, prerequisites, or exclusions. It's a basic functional statement without usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_server_infoB
Get information about the Fooocus MCP server configuration.

Returns:
    Dictionary containing server configuration information
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool 'Returns: Dictionary containing server configuration information,' which provides some output context, but it doesn't describe other behavioral traits such as whether this is a read-only operation, potential rate limits, authentication requirements, or error conditions. For a tool with zero annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: it starts with the core purpose in the first sentence and follows with return information. There's no wasted text or redundancy. However, it could be slightly more structured (e.g., separating purpose and return into distinct sections), but it's efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (simple read operation with no parameters) and the presence of an output schema (which handles return values), the description is somewhat complete. It covers the purpose and mentions the return type, but it lacks behavioral context like safety or usage guidelines. With no annotations and simple schema, the description does the minimum viable job but leaves gaps in transparency and guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and the input schema has 100% description coverage (though empty). The description doesn't need to add parameter information since there are none, so it appropriately focuses on the tool's purpose and output. This meets the baseline of 4 for tools with no parameters, as there's no parameter semantics to explain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Get information about the Fooocus MCP server configuration.' This specifies both the verb ('Get information') and the resource ('Fooocus MCP server configuration'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from sibling tools like 'generate_image' or 'list_available_styles' (which have different purposes), so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any specific contexts, prerequisites, or exclusions for usage. While the purpose is clear, there's no explicit advice on when this tool is appropriate compared to other tools in the server, leaving the agent to infer usage based on the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_available_stylesA
List all available styles for image generation.

Returns:
    Dictionary containing all available styles organized by category
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool lists styles and returns a dictionary organized by category, which is helpful. However, it lacks details on potential side effects, error conditions, or performance aspects (e.g., caching, rate limits), leaving gaps in behavioral understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and well-structured: the first sentence states the purpose, and the second clarifies the return format. Every sentence adds value without redundancy, making it front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (0 parameters, no annotations, but has an output schema), the description is reasonably complete. It explains what the tool does and the return format, and the output schema likely covers return values in detail. However, it could benefit from more behavioral context or usage guidance to be fully comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so the schema fully documents the lack of inputs. The description doesn't add parameter details, which is unnecessary here. A baseline of 4 is appropriate for zero-parameter tools, as no compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'List all available styles for image generation.' It specifies the verb ('List') and resource ('available styles for image generation'), making the function unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'generate_image' or 'get_server_info', which would be needed for a score of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'generate_image' (which might use styles) or 'get_server_info', nor does it specify prerequisites or contextual triggers. This leaves the agent without usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv1.0.0
    • First observedgenerate_image
    • First observedget_server_info
    • First observedlist_available_styles

TDQS

A3.5/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: generate_image creates images, get_server_info provides configuration details, and list_available_styles enumerates style options. The three functions are orthogonal and cannot be confused for one another.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern (generate_image, get_server_info, list_available_styles) with clear, descriptive names that use snake_case uniformly. There are no deviations in naming conventions.

Tool Count3/5

With only 3 tools, the set feels thin for an image generation API. While the core generate_image tool is present, additional operations like image editing, batch generation, or model management might be expected but are missing, making the scope somewhat limited.

Completeness4/5

The toolset covers the essential workflow: checking server info, listing styles, and generating images. However, there are minor gaps such as no ability to modify or delete generated images, adjust advanced generation parameters beyond the basics, or manage generation history, which could limit agent flexibility.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/raihan0824/mcp-fooocus-api'

If you have feedback or need assistance with the MCP directory API, please join our Discord server