Advanced TTS MCP Server
Utilizes FFmpeg for audio conversion, supporting multiple output formats (WAV, MP3, FLAC, OGG) for the synthesized speech
Offers Node.js implementation for running the TTS server, with TypeScript support for type-safe interactions
Uses ONNX runtime for the Kokoro neural voice models, providing high-quality text-to-speech synthesis with multiple voices and emotional expressions
Provides Python API for interacting with the TTS engine, supporting various speech synthesis operations and batch processing
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Advanced TTS MCP Serverread this welcome message in a friendly female voice with excited emotion"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Advanced TTS MCP Server
A high-quality, feature-rich Text-to-Speech MCP server with native TypeScript implementation. Designed for professional applications requiring natural, expressive speech synthesis with advanced controls and zero external dependencies.
⨠Features
šÆ Advanced Voice Control
10 High-Quality Voices - Male and female voices with distinct personalities
Emotion Control - Neutral, happy, excited, calm, serious, casual, confident
Dynamic Pacing - Natural, conversational, presentation, tutorial, narrative modes
Speed & Volume - Precise control from 0.25x to 3.0x speed, 0.1x to 2.0x volume
š Professional Capabilities
Streaming Audio - Real-time synthesis and playback
Batch Processing - Handle multiple text segments efficiently
Multiple Formats - WAV, MP3, FLAC, OGG output support
Natural Speech Enhancement - Automatic pause insertion and emotion markers
Queue Management - Handle multiple concurrent requests
š§ MCP Integration
6 Powerful Tools - Complete synthesis, batch processing, voice management
2 Rich Resources - Voice capabilities and usage examples
Real-time Status - Track processing progress and manage requests
File Management - Save, list, and organize audio outputs
Related MCP server: Fish Audio MCP Server
š Quick Start
Option 1: Deploy to Smithery.ai (Recommended)
šÆ One-Click Deployment to Smithery Platform
Deploy Now: Visit Smithery.ai and import this repository
Configure: Set your preferred voice and speech settings
Use Instantly: Access via Claude Desktop or any MCP-compatible client
Benefits:
ā Zero setup required
ā Automatic scaling and updates
ā No model downloads needed
ā Enterprise-grade hosting
š Full Smithery Deployment Guide ā
Option 2: Local Installation
Prerequisites:
Node.js 18+
Installation:
Clone the repository
git clone https://github.com/samihalawa/advanced-tts-mcp.git
cd advanced-tts-mcpInstall dependencies
npm installConfigure Claude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"advanced-tts": {
"command": "node",
"args": ["dist/index.js"],
"cwd": "/path/to/advanced-tts-mcp"
}
}
}Start using!
# Build TypeScript
npm run build
# Start server
npm startRestart Claude Desktop and start synthesizing with natural, expressive voices.
šļø Available Voices
Voice ID | Name | Gender | Description |
| Heart | Female | Warm, friendly voice (default) |
| Sky | Female | Clear, bright voice |
| Bella | Female | Elegant, sophisticated voice |
| Sarah | Female | Professional, confident voice |
| Nicole | Female | Gentle, soothing voice |
| Adam | Male | Strong, authoritative voice |
| Michael | Male | Friendly, approachable voice |
| Emma | Female | Young, energetic voice |
| Isabella | Female | Mature, expressive voice |
| Lewis | Male | Deep, resonant voice |
š Usage Examples
Basic Synthesis
# Simple text-to-speech
await synthesize_speech(
text="Hello! Welcome to Advanced TTS.",
voice_id="af_heart"
)Emotional Expression
# Excited announcement
await synthesize_speech(
text="This is amazing news! You're going to love this new feature!",
voice_id="af_heart",
emotion="excited",
pacing="conversational",
speed=1.1
)Professional Presentation
# Tutorial narration
await synthesize_speech(
text="Step one: Open your browser. Step two: Navigate to the website.",
voice_id="am_adam",
emotion="calm",
pacing="tutorial",
speed=0.9
)Batch Processing
# Multiple segments with pauses
await batch_synthesize(
segments=[
"Welcome to our presentation.",
"Today we'll cover three main topics.",
"Let's begin with the first topic."
],
voice_id="af_sarah",
emotion="confident",
pacing="presentation",
merge_output=True,
segment_pause=1.0,
save_file=True
)š ļø Available Tools
synthesize_speech
Convert text to natural speech with full control over voice characteristics.
Parameters:
text- Text to synthesize (max 10,000 chars)voice_id- Voice selection (see table above)speed- Speech rate (0.25-3.0)emotion- Voice emotion (neutral, happy, excited, calm, serious, casual, confident)pacing- Speech style (natural, conversational, presentation, tutorial, narrative, fast, slow)volume- Audio volume (0.1-2.0)output_format- File format (wav, mp3, flac, ogg)save_file- Save to file (boolean)filename- Custom filename
batch_synthesize
Process multiple text segments efficiently with optional merging.
Parameters:
segments- List of text segmentsmerge_output- Combine into single filesegment_pause- Pause between segments (0.0-5.0s)All synthesis parameters from above
get_voices
Retrieve complete voice information and capabilities.
get_status
Check processing status for synthesis requests.
cancel_request
Cancel active synthesis operations.
list_output_files
Browse saved audio files with metadata.
šļø Voice Controls
Emotions
Neutral - Standard, professional tone
Happy - Upbeat, cheerful expression
Excited - Enthusiastic, energetic delivery
Calm - Relaxed, soothing tone
Serious - Formal, authoritative delivery
Casual - Relaxed, conversational style
Confident - Assured, professional tone
Pacing Styles
Natural - Balanced, human-like rhythm
Conversational - Casual discussion pace
Presentation - Professional speaking rhythm
Tutorial - Educational, clear delivery
Narrative - Storytelling pace
Fast - Quick delivery (1.2x base speed)
Slow - Deliberate delivery (0.8x base speed)
šµ Audio Formats
Format | Quality | Use Case |
WAV | Uncompressed | Highest quality, editing |
MP3 | Compressed | Web, streaming, sharing |
FLAC | Lossless | Archival, high-quality storage |
OGG | Compressed | Open source alternative |
š§ Configuration
Environment Variables
# Model paths (optional)
KOKORO_MODEL_PATH=./kokoro-v1.0.onnx
KOKORO_VOICES_PATH=./voices-v1.0.bin
# Output settings
TTS_OUTPUT_DIR=./audio_output
TTS_MAX_QUEUE_SIZE=100
# Audio settings
TTS_DEFAULT_VOICE=af_heart
TTS_ENABLE_STREAMING=trueServer Configuration
config = ServerConfig(
model_path="./kokoro-v1.0.onnx",
voices_path="./voices-v1.0.bin",
output_dir="./audio_output",
max_queue_size=100,
enable_streaming=True,
default_voice="af_heart"
)šļø Architecture
āāā src/advanced_tts/
ā āāā __init__.py # Package initialization
ā āāā server.py # MCP server implementation
ā āāā engine.py # Kokoro TTS engine wrapper
ā āāā models.py # Data models and validation
ā āāā utils.py # Utility functions
āāā pyproject.toml # Project configuration
āāā README.md # Documentation
āāā LICENSE # MIT Licenseš¤ Contributing
Contributions welcome! Areas for improvement:
Additional voice models
Real-time streaming synthesis
Advanced audio effects
Multi-language support
Performance optimizations
š License
MIT License - see LICENSE for details.
š Acknowledgments
Kokoro TTS - High-quality neural voice synthesis
MCP Protocol - Seamless AI model integration
FastMCP - Efficient server framework
Developed by Sami Halawa
Transform your text into natural, expressive speech with Advanced TTS MCP Server.
Available Tools
5 toolsbatch_synthesizeC
Synthesize multiple text segments with optional merging and intelligent pacing
| Name | Required | Description | Default |
|---|---|---|---|
| emotion | No | Voice emotion for all segments | neutral |
| filename | No | Custom filename for saved audio | |
| mergeOutput | No | Merge segments into single file | |
| outputFormat | No | Audio output format | wav |
| pacing | No | Speech pacing style for all segments | natural |
| saveFile | No | Save audio to file | |
| segmentPause | No | Pause between segments in seconds | |
| segments | Yes | List of text segments to synthesize | |
| speed | No | Speech speed (0.25-3.0) | |
| voiceId | No | Voice to use for all segments | af_heart |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'optional merging and intelligent pacing,' which adds some context about output behavior, but fails to cover critical aspects: whether synthesis is resource-intensive, if there are rate limits, authentication needs, error handling, or what the output entails (e.g., audio files, metadata). For a tool with 10 parameters and no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Synthesize multiple text segments with optional merging and intelligent pacing.' It is front-loaded with the core action and key features, with zero wasted words. Every element earns its place by conveying essential information without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (10 parameters, no output schema, no annotations), the description is incomplete. It lacks details on output behavior (e.g., what is returned, file handling), error conditions, performance implications, and how it differs from siblings like 'synthesize_speech.' For a batch synthesis tool with rich parameters, the description should provide more context to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are documented in the input schema. The description adds minimal value beyond the schemaāit implies batch processing ('multiple text segments') and hints at 'merging' (related to 'mergeOutput') and 'pacing' (related to 'pacing'), but does not elaborate on parameter interactions or semantics. With high schema coverage, the baseline score of 3 is appropriate, as the description does not compensate with additional insights.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Synthesize multiple text segments with optional merging and intelligent pacing.' It specifies the verb ('synthesize'), resource ('multiple text segments'), and key optional features ('merging' and 'intelligent pacing'), making the intent unambiguous. However, it does not explicitly differentiate from sibling tools like 'synthesize_speech'ālikely a batch version versus single synthesisāwhich prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions 'optional merging and intelligent pacing,' which hints at features, but does not specify scenarios, prerequisites, or comparisons to sibling tools (e.g., 'synthesize_speech' for single segments). Without explicit when-to-use or when-not-to-use advice, the agent lacks context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statusC
Get processing status for a synthesis request
| Name | Required | Description | Default |
|---|---|---|---|
| requestId | Yes | Request ID to check status for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool retrieves status but doesn't describe what the status includes (e.g., pending, completed, failed), whether it's a read-only operation, potential errors (e.g., invalid request ID), or rate limits. This leaves significant gaps for a tool that likely interacts with asynchronous processes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without unnecessary words. Every part of the sentence ('Get processing status for a synthesis request') contributes directly to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete for a status-checking tool. It doesn't explain what information is returned (e.g., status states, progress percentages, error messages) or behavioral aspects like idempotency or polling requirements. This is inadequate for guiding an agent in a synthesis workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the single parameter 'requestId' clearly documented in the schema. The description adds no additional meaning beyond implying the parameter is for a synthesis request, which is already evident from the tool's context. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('processing status for a synthesis request'), making the purpose immediately understandable. It distinguishes this from sibling tools like 'synthesize_speech' (which creates requests) and 'list_output_files' (which lists results). However, it doesn't explicitly differentiate from 'batch_synthesize' or 'get_voices', which are related but serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a request ID from a previous synthesis operation), exclusions, or comparisons to siblings like 'list_output_files' for retrieving results. Usage is implied but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_voicesB
Get list of available voices with their capabilities and supported features
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions what information is returned but doesn't address important behavioral aspects like whether this is a read-only operation, if there are rate limits, authentication requirements, or what format the response takes. The description provides basic output content but lacks operational context needed for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the essential information without any wasted words. It's front-loaded with the core purpose and adds specific detail about what's included in the response. Every word earns its place in this compact description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read operation with no output schema, the description provides adequate but minimal information. It tells what the tool does and what information it returns, but doesn't address format, structure, or behavioral constraints. Given the lack of annotations and output schema, more detail about response format or operational considerations would improve completeness for this type of tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema already fully documents the parameter situation. The description appropriately doesn't discuss parameters since none exist, and it focuses instead on what the tool returns. This meets the baseline expectation for parameterless tools while adding value about the return content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Get') and resource ('list of available voices'), making the purpose immediately understandable. It adds specificity about what information is returned ('capabilities and supported features'), which goes beyond just listing voices. However, it doesn't explicitly differentiate from sibling tools like 'list_output_files' or 'get_status', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'synthesize_speech' or 'batch_synthesize'. There's no mention of prerequisites, typical use cases, or when this tool would be appropriate versus when other tools might be better suited. The agent must infer usage context from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_output_filesB
List saved audio files in the output directory with metadata
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but does not disclose any behavioral traits such as whether it requires specific permissions, how it handles errors, if it has rate limits, or what the output format looks like. This is a significant gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that directly states the tool's purpose without any waste. It is front-loaded with the core action and resource, making it highly concise and easy to parse, earning its place with no extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It does not explain what metadata is included, how files are sorted or filtered, or what the return values look like. For a tool that lists files with metadata, more context is needed to fully understand its behavior and output, making it inadequate for the complexity involved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so there are no parameters to document. The description does not need to add parameter semantics beyond the schema, and it appropriately avoids unnecessary details. A baseline of 4 is applied as it handles the zero-parameter case efficiently without redundancy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('saved audio files in the output directory with metadata'), making the purpose specific and understandable. However, it does not explicitly differentiate from sibling tools like 'get_status' or 'get_voices', which might also involve listing or retrieving information, so it misses full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any context, prerequisites, or exclusions, such as when to prefer 'list_output_files' over 'get_status' for checking file availability or other sibling tools. This lack of usage context leaves the agent without clear direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
synthesize_speechC
Convert text to speech with advanced voice controls and natural expression
| Name | Required | Description | Default |
|---|---|---|---|
| emotion | No | Voice emotion | neutral |
| filename | No | Custom filename for saved audio | |
| outputFormat | No | Audio output format | wav |
| pacing | No | Speech pacing style | natural |
| saveFile | No | Save audio to file | |
| speed | No | Speech speed (0.25-3.0) | |
| text | Yes | Text to convert to speech | |
| voiceId | No | Voice to use for synthesis | af_heart |
| volume | No | Audio volume (0.1-2.0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While 'Convert text to speech' implies a creation/write operation, it doesn't address key behavioral aspects: whether this is a synchronous or asynchronous process, potential rate limits, authentication requirements, file storage implications when saveFile is true, or what happens on failure. The mention of 'advanced voice controls' is vague and doesn't provide concrete behavioral information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that gets straight to the point. 'Convert text to speech' is front-loaded with the core function, followed by additional context. There's no wasted verbiage or redundancy. However, it could be slightly more structured by separating core function from additional capabilities.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, no annotations, and no output schema, the description is insufficiently complete. It doesn't explain what the tool returns (audio data? file path? success status?), doesn't address error conditions, and provides minimal behavioral context. The combination of complex parameters and lack of structured metadata requires a more comprehensive description to guide proper tool usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal value beyond the input schema, which has 100% coverage. 'with advanced voice controls and natural expression' vaguely references the emotion, pacing, and voiceId parameters but doesn't provide additional semantic context. The schema already comprehensively documents all 9 parameters with descriptions, defaults, enums, and constraints, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Convert text to speech' specifies the verb and resource. It adds 'with advanced voice controls and natural expression' which provides additional context about capabilities. However, it doesn't explicitly differentiate from sibling tools like batch_synthesize or get_voices, which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention batch_synthesize for multiple texts, get_voices for voice selection, or list_output_files for file management. There's no context about prerequisites, limitations, or appropriate use cases beyond the basic function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
5 tool updates
v1.0.0- First observed
batch_synthesize - First observed
get_status - First observed
get_voices - First observed
list_output_files - First observed
synthesize_speech
TDQS
Each tool has a distinct, non-overlapping purpose: batch_synthesize handles multiple segments, synthesize_speech handles single conversions, get_status checks request status, get_voices lists voice options, and list_output_files manages saved files. The descriptions clearly differentiate their functions, eliminating ambiguity.
All tools follow a consistent verb_noun naming pattern (e.g., batch_synthesize, get_status, get_voices, list_output_files, synthesize_speech). The verbs (batch_, get_, list_, synthesize_) are appropriate and uniform, making the set predictable and easy to understand.
With 5 tools, the server is well-scoped for a TTS (text-to-speech) domain. The count is appropriate, covering core operations like synthesis, status checking, voice management, and file listing without being too sparse or bloated. Each tool earns its place in the workflow.
The tool set covers essential TTS operations: synthesis (single and batch), status tracking, voice discovery, and output management. A minor gap exists in lacking explicit tools for deleting or managing output files beyond listing, but agents can likely work around this, and core workflows are well-supported.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Hosted pay-per-use TTS: 54 neural voices, 9 languages incl. Brazilian Portuguese. $10 free credits.
Pronunciation assessment, phoneme scoring, speaker voice ID, audio transcription, speech synthesis.
Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to convert text to speech using Microsoft Edge's Text-to-Speech service with customizable voice options, speech rate, volume, and pitch parameters.MIT
- AlicenseBqualityDmaintenanceEnables natural language-driven speech synthesis using Fish Audio's Text-to-Speech API, supporting multiple voices, streaming, and flexible configuration.221MIT

leanvox-mcpofficial
AlicenseNot gradedqualityDmaintenanceEnables text-to-speech generation, voice cloning, dialogue creation, and other TTS operations through natural language in MCP-compatible AI assistants.15MIT- AlicenseAqualityAmaintenanceGive your AI agent a voice with x402 pay-per-call speech synthesis, offering 20 voices, 10 personas, 31 languages, and granular controls.46198MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/samihalawa/advanced-tts-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server