MCP Fooocus API
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Fooocus APIcreate a fantasy landscape with dragons and castles at sunset"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Fooocus API
A Model Context Protocol (MCP) server that provides text-to-image generation capabilities through the Fooocus Stable Diffusion API.
Features
Text-to-Image Generation: Generate high-quality images from text prompts
Intelligent Style Selection: Automatically selects 1-3 appropriate styles based on your prompt
Custom Style Override: Manually specify styles from 300+ available options
Multiple Performance Modes: Choose between Speed, Quality, and Extreme Speed
Configurable Aspect Ratios: Support for various image dimensions
Environment-based Configuration: Easy API endpoint configuration via
.envfile
Related MCP server: Volcengine Image Generation MCP Server
Installation
Using uv (Recommended)
Install directly from GitHub:
uv add git+https://github.com/raihan0824/mcp-fooocus-api.gitOr install from PyPI (when published):
uv add mcp-fooocus-apiRun with uvx:
uvx --from git+https://github.com/raihan0824/mcp-fooocus-api.git mcp-fooocus-apiUsing pip
pip install git+https://github.com/raihan0824/mcp-fooocus-api.gitOr from PyPI (when published):
pip install mcp-fooocus-apiDevelopment Installation
# Clone the repository
git clone https://github.com/raihan0824/mcp-fooocus-api.git
cd mcp-fooocus-api
# Install with uv
uv sync --dev
# Or install with pip
pip install -e ".[dev]"Configuration
Copy the example environment file:
cp .env.example .envEdit the
.envfile to configure your Fooocus API endpoint:
FOOOCUS_API_URL=http://103.125.100.56:8888/v1/generation/text-to-imageUsage
Available Tools
The MCP server provides three main tools:
1. generate_image
Generate an image using the Fooocus API.
Parameters:
prompt(required): Text description of the image to generateperformance(optional): Performance setting - "Speed" (default), "Quality", or "Extreme Speed"custom_styles(optional): Comma-separated list of custom stylesaspect_ratio(optional): Image dimensions (default: "1024*1024")
Example:
{
"prompt": "A serene landscape with mountains and a lake at sunset",
"performance": "Quality",
"aspect_ratio": "1024*1024"
}2. list_available_styles
Lists all available styles organized by category.
Returns:
Total number of available styles
Styles organized by categories (Fooocus, SAI, MRE, Art Styles, etc.)
Available performance options
3. get_server_info
Get information about the server configuration and capabilities.
Returns:
Server version and name
Configured API endpoint
Available features
Performance options
Style Categories
The server includes 300+ styles organized into categories:
Fooocus Styles: Native Fooocus styles (V2, Enhance, Sharp, etc.)
SAI Styles: Stability AI styles (Photographic, Digital Art, Anime, etc.)
Art Styles: Classical art movements (Renaissance, Impressionist, Cubist, etc.)
Photography: Various photography styles (Film Noir, HDR, Macro, etc.)
Game Styles: Video game-inspired styles (Minecraft, Pokemon, Retro, etc.)
Futuristic: Sci-fi and cyberpunk styles
And many more...
Intelligent Style Selection
When you don't specify custom styles, the server automatically selects appropriate styles based on your prompt:
"renaissance portrait" → Selects "Artstyle Renaissance"
"cyberpunk city" → Selects "Futuristic Cyberpunk Cityscape"
"anime character" → Selects "SAI Anime"
"realistic photo" → Selects "SAI Photographic"
"watercolor painting" → Selects "Artstyle Watercolor"
Running the Server
As an MCP Server
Add to your MCP client configuration:
{
"mcpServers": {
"fooocus": {
"command": "uvx",
"args": ["--from", "git+https://github.com/raihan0824/mcp-fooocus-api.git", "mcp-fooocus-api"]
}
}
}Or if installed from PyPI:
{
"mcpServers": {
"fooocus": {
"command": "uvx",
"args": ["mcp-fooocus-api"]
}
}
}Standalone Server
You can also run the server directly:
# With uv
uvx --from git+https://github.com/raihan0824/mcp-fooocus-api.git mcp-fooocus-api --port 3000 --host localhost
# Or if installed locally
python -m mcp_fooocus_api.server --port 3000 --host localhostAPI Response Format
Successful generation returns:
{
"success": true,
"prompt": "Your prompt here",
"selected_styles": ["Style1", "Style2"],
"performance": "Speed",
"aspect_ratio": "1024*1024",
"result": {
// Fooocus API response data
}
}Error responses include:
{
"success": false,
"error": "Error description",
"prompt": "Your prompt here",
"selected_styles": ["Style1", "Style2"]
}Requirements
Python 3.8+
Access to a Fooocus API endpoint
Internet connection for API requests
Dependencies
mcp>= 1.0.0httpx>= 0.27python-dotenv>= 1.0.0pydantic>= 2.7.2, < 3.0.0
Development
To set up for development:
Clone the repository
Install dependencies:
pip install -e .Configure your
.envfileRun the server:
python -m mcp_fooocus_api.server
License
MIT License - see LICENSE file for details.
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Support
For issues and questions, please visit the GitHub repository.
Available Tools
3 toolsgenerate_imageC
Generate an image using Fooocus Stable Diffusion API.
Args:
prompt: Text description of the image to generate
performance: Performance setting (Speed, Quality, Extreme Speed)
custom_styles: Comma-separated list of custom styles (optional, will auto-select if not provided)
aspect_ratio: Image aspect ratio (default: 1024*1024)
Returns:
Dictionary containing generation result with image URLs
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| performance | No | Speed | |
| custom_styles | No | ||
| aspect_ratio | No | 1024*1024 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the API and return format but lacks critical behavioral details: no information on rate limits, authentication needs, error handling, or what happens if generation fails. For a generative tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with sections for Args and Returns, making it easy to scan. It's concise with no wasted sentences, though it could be more front-loaded by emphasizing the core purpose first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (image generation with 4 parameters) and no annotations, the description is moderately complete. It covers parameters and return format, but lacks behavioral context and usage guidelines. The output schema exists, so describing return values isn't needed, but other gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It lists all 4 parameters with brief explanations (e.g., 'Text description of the image to generate' for prompt), adding meaning beyond the schema's titles. However, it doesn't fully detail constraints like valid aspect ratio formats or performance options, leaving gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Generate an image using Fooocus Stable Diffusion API.' It specifies the verb ('Generate') and resource ('image'), though it doesn't explicitly differentiate from sibling tools like list_available_styles. The purpose is specific but lacks sibling comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description mentions the API but doesn't specify scenarios, prerequisites, or exclusions. It's a basic functional statement without usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_server_infoB
Get information about the Fooocus MCP server configuration.
Returns:
Dictionary containing server configuration information
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool 'Returns: Dictionary containing server configuration information,' which provides some output context, but it doesn't describe other behavioral traits such as whether this is a read-only operation, potential rate limits, authentication requirements, or error conditions. For a tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded: it starts with the core purpose in the first sentence and follows with return information. There's no wasted text or redundancy. However, it could be slightly more structured (e.g., separating purpose and return into distinct sections), but it's efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (simple read operation with no parameters) and the presence of an output schema (which handles return values), the description is somewhat complete. It covers the purpose and mentions the return type, but it lacks behavioral context like safety or usage guidelines. With no annotations and simple schema, the description does the minimum viable job but leaves gaps in transparency and guidelines.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and the input schema has 100% description coverage (though empty). The description doesn't need to add parameter information since there are none, so it appropriately focuses on the tool's purpose and output. This meets the baseline of 4 for tools with no parameters, as there's no parameter semantics to explain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get information about the Fooocus MCP server configuration.' This specifies both the verb ('Get information') and the resource ('Fooocus MCP server configuration'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from sibling tools like 'generate_image' or 'list_available_styles' (which have different purposes), so it doesn't reach the highest score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any specific contexts, prerequisites, or exclusions for usage. While the purpose is clear, there's no explicit advice on when this tool is appropriate compared to other tools in the server, leaving the agent to infer usage based on the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_stylesA
List all available styles for image generation.
Returns:
Dictionary containing all available styles organized by category
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool lists styles and returns a dictionary organized by category, which is helpful. However, it lacks details on potential side effects, error conditions, or performance aspects (e.g., caching, rate limits), leaving gaps in behavioral understanding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and well-structured: the first sentence states the purpose, and the second clarifies the return format. Every sentence adds value without redundancy, making it front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (0 parameters, no annotations, but has an output schema), the description is reasonably complete. It explains what the tool does and the return format, and the output schema likely covers return values in detail. However, it could benefit from more behavioral context or usage guidance to be fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so the schema fully documents the lack of inputs. The description doesn't add parameter details, which is unnecessary here. A baseline of 4 is appropriate for zero-parameter tools, as no compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'List all available styles for image generation.' It specifies the verb ('List') and resource ('available styles for image generation'), making the function unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'generate_image' or 'get_server_info', which would be needed for a score of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'generate_image' (which might use styles) or 'get_server_info', nor does it specify prerequisites or contextual triggers. This leaves the agent without usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v1.0.0- First observed
generate_image - First observed
get_server_info - First observed
list_available_styles
TDQS
Each tool has a clearly distinct purpose with no overlap: generate_image creates images, get_server_info provides configuration details, and list_available_styles enumerates style options. The three functions are orthogonal and cannot be confused for one another.
All tools follow a consistent verb_noun pattern (generate_image, get_server_info, list_available_styles) with clear, descriptive names that use snake_case uniformly. There are no deviations in naming conventions.
With only 3 tools, the set feels thin for an image generation API. While the core generate_image tool is present, additional operations like image editing, batch generation, or model management might be expected but are missing, making the scope somewhat limited.
The toolset covers the essential workflow: checking server info, listing styles, and generating images. However, there are minor gaps such as no ability to modify or delete generated images, adjust advanced generation parameters beyond the basics, or manage generation history, which could limit agent flexibility.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Generate images, video, music, voice and 3D through one API. 30 tools, 200+ models.
Image, video, music and text generation across 100+ models through one endpoint.
Generate reproducible image, video, and audio assets with leading models and your own provider keys.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables text-to-image generation, image editing, and multi-image composition using Google's Gemini 2.5 Flash Image API. Supports flexible aspect ratios and character consistency across generations.1-
- FlicenseBqualityDmaintenanceEnables AI-powered text-to-image generation using Volcengine's API with support for multiple image sizes, customizable parameters like guidance scale and seed, and flexible output formats.1-
- FlicenseBqualityDmaintenanceEnables text-to-image generation using the fal.ai GPT image-1 API. It provides a tool to generate images with customizable parameters like size, quality, and background settings via natural language prompts.1-
- AlicenseNot gradedqualityCmaintenanceEnables LLM applications to generate, edit, describe, upscale, remix, reframe, and replace backgrounds in images using the Ideogram AI API.20MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/raihan0824/mcp-fooocus-api'
If you have feedback or need assistance with the MCP directory API, please join our Discord server