Audio Processing
Services for manipulating, generating, and working with audio content. Includes audio synthesis, processing, playback control, and format conversion capabilities.
MCP ServersBrowse all →
AlicenseAqualityAmaintenanceVoice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.649MIT
GlianaAI MCP Serverofficial
AlicenseAqualityBmaintenanceEnables pay-per-call access to 90+ generative AI models and utility tools via any MCP client, with no signup or API key, using wallet-based USDC payments on Base, Tempo, or Solana.7291MIT- MIT

@paxalabs/mcpofficial
AlicenseAqualityBmaintenanceEnables agents to speak Thai and English audio through local speakers, manage a playback queue, save speech to files, translate text into Thai, and OCR PDFs and images.10MIT- AlicenseAqualityCmaintenanceRemove vocals, extract instrumentals, and split any song into up to six stems — directly from Claude Desktop, Cursor, or any MCP client. Supports local audio files, YouTube URLs, and SoundCloud track1120MIT

RunAPI MCP Serverofficial
AlicenseAqualityBmaintenanceConnects MCP-compatible coding tools to RunAPI for AI image, video, music, text-to-speech, and LLM generation using 130+ models from leading providers.938755Apache 2.0
Clipia MCPofficial
AlicenseAqualityAmaintenanceGenerate AI images, video, speech, and music from Claude, ChatGPT, Cursor, and other MCP clients through the Clipia API.15MIT
MMAudio MCPofficial
AlicenseBqualityCmaintenanceEnables AI-powered video-to-audio and text-to-audio generation using MMAudio's API. Create synchronized audio from video content or generate audio from text descriptions with configurable parameters.3173MIT- AlicenseAqualityAmaintenanceRun AI workflows hosted on Glif.app via MCP, including ComfyUI-based image generators, meme generators, selfies, chained LLM calls, and more6383203MIT
- AlicenseAqualityDmaintenanceOfficial AllVoiceLab Model Context Protocol (MCP) server, supporting interaction with powerful text-to-speech and video translation APIs.1259MIT

@speechweave/mcpofficial
AlicenseAqualityBmaintenanceMCP server for SpeechWeave transcription, enabling AI assistants to transcribe local files and URLs via wait-first or async tools.6103MIT
ElevenLabs MCP Serverofficial
AlicenseAqualityFmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,536MIT- AlicenseAqualityAmaintenanceText to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.4MIT

Sonilo MCPofficial
AlicenseAqualityBmaintenanceAn MCP (Model Context Protocol) server that exposes Sonilo's licensed music and sound-effects API to MCP-compatible clients (Claude Code, Claude Desktop, Codex).954MIT- AlicenseAqualityAmaintenanceGive your AI agent a voice with x402 pay-per-call speech synthesis, offering 20 voices, 10 personas, 31 languages, and granular controls.46198MIT

ZeroTrue MCP Serverofficial
AlicenseAqualityCmaintenanceEnables detection of AI-generated content in text, images, video, and audio via the ZeroTrue API, supporting multiple analysis tools and MCP-compatible clients.617MIT
Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.225MIT- AlicenseBqualityCmaintenanceAn AI composition partner that remote-controls Steinberg Dorico through the Model Context Protocol, enabling composers to create, edit, play back, and refine musical scores using natural language.6291AGPL 3.0

mocoVoice MCP Serverofficial
AlicenseAqualityCmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.63MIT- AlicenseAqualityCmaintenanceGaudio Lab Audio AI — Stem Separation, DME Separation, AI Text Sync7811MIT
- AlicenseAqualityBmaintenanceEnables AI agents to analyze local or URL-hosted audio files with TrackTag, returning professional music metadata such as BPM, key, genres, moods, and 35+ fields. Runs locally on your machine, using your TrackTag API key and credit balance for analyses.5461MIT
- AlicenseAqualityCmaintenanceEnables text-to-speech synthesis through a streaming, GPU-accelerated gateway, exposing a single tool that returns playable WAV audio with configurable voice and speed.1MIT
- AlicenseAqualityCmaintenanceA job queue-based FFmpeg wrapper enabling AI assistants to perform video processing tasks such as trimming, format conversion, resolution change, and subtitle conversion through natural language.8MIT
- AlicenseAqualityBmaintenanceEnables local analysis of PCM WAV files to produce bounded audio observations such as activity segments, clipping indicators, duration, sample rate, peak, and RMS level without uploading recordings.1MIT
- AlicenseAqualityDmaintenanceAnalyzes speech audio to detect emotions, urgency, and sarcasm using prosodic features.51MIT
- AlicenseBqualityAmaintenanceConnects Ableton Live to Claude AI through the Model Context Protocol, enabling AI-assisted music production by allowing Claude to directly interact with and control Ableton Live sessions.162,950MIT
- AlicenseAqualityAmaintenanceAn interactive digital audio workstation as an MCP server, enabling music production with a channel rack, piano roll, mixer, effects, automation, microphone recording, and offline WAV rendering.2661MIT
- AlicenseAqualityCmaintenanceAI-powered speech tools by Brainiall: pronunciation assessment with phoneme-level feedback, speech-to-text with language detection, and text-to-speech with multiple voices.41MIT
- AlicenseAqualityNot gradedmaintenanceEnables interaction with MiniMax AI APIs for text-to-speech, voice cloning, video generation, image generation, and music creation through MCP clients like Claude Desktop and Cursor.9-
- AlicenseAqualityBmaintenanceEnables reading and controlling a Universal Audio Apollo interface: channel names, faders, preamp gain, cue sends, monitor control, and more via the UA Mixer Engine's local TCP API.21MIT
MCP ConnectorsBrowse all →
Generate highly realistic Text to Speech voiceovers.
Create and edit images, videos, and audio through Magic Hour's hosted Streamable HTTP MCP server.
Magic Hour MCP lets AI agents create and edit images, videos, and audio using Magic Hour’s hosted generation tools.
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Generate marketing images, videos and audios for campaigns, product content, and brand assets.
Deterministic music theory for agents: analyze, voice, reharmonize, conduct — computed, not guessed
Generate game assets with AI: sprites, 3D models, animations, sound effects, music, and voices.
Generate Suno AI music (v5.5) from any MCP client. Async; billed only on success.
One tool surface for music, image, video, and audio generation across Suno, Grok Imagine, Seedance, Kling, Hailuo, Wan, VEO, Ideogram, and GPT Image 2. Generate, edit, upscale, reframe, and master through one credit pool. Connect in one click with OAuth, no API key required.
Find & cut horizontal and vertical video clips (Shorts/Reels), transcribe & summarize. Pay per job.
Deepfake detection, media intelligence, and invisible watermarking for audio, image, and video via the Resemble AI API, plus docs tools. Remote MCP server (Streamable HTTP) — also published in the official MCP registry as io.github.resemble-ai/resemble-mcp.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Search KYMA369's public frequency library and open sets in the free KYMA app.
Create, co-edit, analyze, publish, and export collaborative step-sequencer sessions through MCP.
Generate text, images, speech, music, and video with any AI model, from one credit balance.
Human-made production music for sync — search by brief or reference, preview, score to picture.
File conversion: PDF, DOCX, STT, TTS, watermarking
Process video, audio, images, and documents with 86+ cloud media processing robots.
Anonymous public tools for Zachary Roth Music. See the published agent boundary before use.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.