Zonos TTS MCP Server
The Zonos TTS MCP Server enables text-to-speech functionality in Claude through the speak_response tool, allowing it to generate and play spoken audio from text with the following capabilities:
Multi-Language Support: Generate speech in different languages (default:
en-us)Emotion Control: Customize the emotional tone (neutral, happy, sad, angry)
PulseAudio Integration: Ensures proper audio playback
MCP Integration: Seamlessly works with Claude's Model Context Protocol
The server requires a running instance of Zonos API with proper Node.js and PulseAudio setup.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Zonos TTS MCP Serverspeak this in French with a happy tone: Bonjour, comment allez-vous?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Zonos MCP Integration
A Model Context Protocol integration for Zonos TTS, allowing Claude to generate speech directly.
Setup
Installing via Smithery
To install Zonos TTS Integration for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @PhialsBasement/zonos-tts-mcp --client claudeManual installation
Make sure you have Zonos running with our API implementation (PhialsBasement/zonos-api)
Install dependencies:
npm install @modelcontextprotocol/sdk axiosConfigure PulseAudio access:
# Your pulse audio should be properly configured for audio playback
# The MCP server will automatically try to connect to your pulse serverBuild the MCP server:
npm run build
# This will create the dist folder with the compiled serverAdd to Claude's config file: Edit your Claude config file (usually in
~/.config/claude/config.json) and add this to themcpServerssection:
"zonos-tts": {
"command": "node",
"args": [
"/path/to/your/zonos-mcp/dist/server.js"
]
}Replace /path/to/your/zonos-mcp with the actual path where you installed the MCP server.
Related MCP server: ElevenLabs MCP Server
Using with Claude
Once configured, Claude automatically knows how to use the speak_response tool:
speak_response(
text="Your text here",
language="en-us", # optional, defaults to en-us
emotion="happy" # optional: "neutral", "happy", "sad", "angry"
)Features
Text-to-speech through Claude
Multiple emotions support
Multi-language support
Proper audio playback through PulseAudio
Requirements
Node.js
PulseAudio setup
Running instance of Zonos API (PhialsBasement/zonos-api)
Working audio output device
Notes
Make sure both the Zonos API server and this MCP server are running
Audio playback requires proper PulseAudio configuration
Available Tools
1 toolspeak_responseD
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| language | No | en-us | |
| emotion | No | neutral |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
- First observed
speak_response
TDQS
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool 'speak_response' stands alone with no other tools to confuse it with, making disambiguation perfect.
Since there is only one tool, naming consistency is inherently perfect. There are no other tools to compare against, so no inconsistencies can exist in the naming pattern.
A single tool for a TTS (Text-to-Speech) server feels thin and under-scoped. TTS typically involves operations like generating, streaming, or managing speech, but this server only offers one tool, which is insufficient for a complete TTS domain coverage.
The server is severely incomplete for a TTS domain. With only a 'speak_response' tool and no description, it lacks essential operations such as configuring voices, controlling speech parameters, or handling different input formats, making it inadequate for typical TTS workflows.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Claude makes real phone calls for you — in many languages, with transcript and outcome back in chat.
Transcribe audio and video with Speechmatics speech-to-text from Claude and any MCP client.
ElevenLabs in natural language: generate speech in any language, create and manage voices, compose m
Use AI models for chat, image, and video generation from Claude Code and other MCP hosts.
Related MCP Servers
- AlicenseDqualityCmaintenanceProvides text-to-speech capabilities through the Model Context Protocol, allowing applications to easily integrate speech synthesis with customizable voices, adjustable speech speed, and cross-platform audio playback support.110MIT
- AlicenseNot gradedqualityDmaintenanceOfficial ElevenLabs Model Context Protocol server that enables AI assistants like Claude to interact with Text to Speech and audio processing APIs, allowing them to generate speech, clone voices, transcribe audio, and create soundscapes.1MIT
- AlicenseNot gradedqualityDmaintenanceEnables voice-based interactions with Claude by converting text to speech using Kokoro TTS and transcribing user responses using NVIDIA NeMo ASR, creating interactive voice dialogues.MIT
- AlicenseNot gradedqualityDmaintenanceProvides text-to-speech generation using the Kokoro-82M model, enabling AI assistants to generate voiceovers and audio content directly within Claude Desktop and Cursor.14Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/PhialsBasement/Zonos-TTS-MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server