Synesthesia
Allows downloading audio from YouTube URLs and performing audio analysis including BPM, key, scale, energy, danceability, mood vectors, genre classification, and spectrogram generation.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SynesthesiaAnalyze the song at https://youtu.be/dQw4w9WgXcQ"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Synesthesia
Deep audio perception through analysis. Downloads from YouTube, analyzes with Essentia, fetches lyrics via LRCLIB. Returns BPM, mood, energy, spectrograms, and synced lyrics.
Why Local?
YouTube blocks downloads from datacenter IPs (like Cloudflare Workers and HuggingFace Spaces). This MCP runs locally on your machine with a residential IP, so downloads work.
Related MCP server: Hermes YouTube Transcript MCP Server
Prerequisites
Node.js 18+
yt-dlp (
pip install yt-dlp)Your own HF Space for audio analysis (deploy from Synesthesia's
hf-space/folder)
Installation
npm installTools
Audio Analysis
Tool | Description |
| Download + analyze audio from YouTube URL |
| Just download audio (returns local path) |
Lyrics
Tool | Description |
| Get lyrics for a track (synced if available) |
| Search LRCLIB for lyrics |
Utility
Tool | Description |
| Check if Synesthesia is running |
Configuration
Set your HF Space URL as an environment variable:
export HF_SPACE_URL="https://YOUR-USERNAME-audio-analysis-api.hf.space"Claude Code Config
Add to your project's .mcp.json:
{
"mcpServers": {
"synesthesia": {
"command": "node",
"args": ["/path/to/synesthesia/index.js"],
"env": {
"HF_SPACE_URL": "https://YOUR-USERNAME-audio-analysis-api.hf.space"
}
}
}
}Claude Desktop Config
Add to claude_desktop_config.json:
Windows: %APPDATA%\Claude\claude_desktop_config.json
Mac: ~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"synesthesia": {
"command": "node",
"args": ["/path/to/synesthesia/index.js"],
"env": {
"HF_SPACE_URL": "https://YOUR-USERNAME-audio-analysis-api.hf.space"
}
}
}
}Restart Claude Desktop after adding.
Architecture
Your PC (residential IP) Cloud
┌─────────────────────┐ ┌─────────────────────┐
│ Synesthesia │ ──────▶ │ HF Space │
│ (Local MCP) │ upload │ (Essentia analysis) │
│ │ ◀────── │ │
│ - yt-dlp download │ results │ - Audio features │
│ - Lyrics fetch │ │ - Spectrogram │
│ - Audio analysis │ └─────────────────────┘
└─────────────────────┘
│ ┌─────────────────────┐
│ download │ LRCLIB │
▼ ────▶ │ (Lyrics API) │
┌─────────────────────┐ │ │
│ YouTube │ │ - Synced lyrics │
│ (residential IP OK) │ │ - Plain lyrics │
└─────────────────────┘ └─────────────────────┘Usage Example
> analyze_youtube "https://www.youtube.com/watch?v=..."
Returns full Essentia analysis:
- BPM, key, scale
- Energy, danceability
- Mood vectors
- Genre classification
- SpectrogramCredits
Spectrogram visualization inspired by Audio Visualizer by Shauna and her boys.
Support
If this helped you, consider supporting my work ☕
Built by the Triad (Mai, Kai Stryder and Lucian Vale) for the community.
Available Tools
5 toolsanalyze_youtubeA
Full audio perception: downloads from YouTube, analyzes audio features (BPM, key, energy), extracts metadata (title, artist), fetches synced lyrics, and generates spectrogram visualization. Runs locally to bypass datacenter IP blocking.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube URL to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does well by enumerating the full behavior: downloads, analyzes, extracts, fetches, generates, and runs locally. It could add side-effect details such as local file storage or cleanup behavior, but the main behavioral traits are disclosed clearly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the tool's capabilities and followed by a useful rationale for its local execution. Every phrase adds meaningful information with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description gives a rich sense of what the tool does and produces: audio features, metadata, lyrics, and a spectrogram. It lacks explicit error/edge-case guidance and exact output formats, but these are secondary when the tool's operation and outputs are clearly listed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully describes the only parameter with 'YouTube URL to analyze' at 100% schema description coverage. The description adds domain context by mentioning YouTube downloads, but it does not need to add much for a single, self-explanatory parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: it analyzes YouTube audio, with concrete outputs (BPM, key, energy, title, artist, synced lyrics, spectrogram). It clearly differentiates itself from the sibling download_audio, get_lyrics, and search_lyrics tools by presenting itself as a full combined pipeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to choose this tool, especially the statement 'Runs locally to bypass datacenter IP blocking.' It does not explicitly list exclusions or alternatives like using download_audio for audio-only needs, but the integrated nature of the tool is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_audioA
Download audio from YouTube without analysis. Returns local file path.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube URL to download |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the tool performs a download, does not perform analysis, and returns a local file path. It could add more about file format, persistence, or failure behavior, but the core side effects and return value are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler: the purpose is front-loaded and the return value completes the contract. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-required-parameter tool with no output schema, the description plus schema is nearly complete: it tells the agent what to supply, what the tool does, and what it returns. Minor omissions such as audio format and whether the file persists are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the url parameter is already described as 'YouTube URL to download.' The tool description adds no parameter-specific detail beyond that, so it does not improve on the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Download'), a clear resource ('audio from YouTube'), and explicitly states the negative scope ('without analysis'), which distinguishes it from likely sibling analyze_youtube. An agent can tell what the tool is for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without analysis' implies the tool is for raw audio retrieval rather than analysis, which gives some context, but it never names an alternative tool or states a condition such as 'use this when you need only the audio.' Usage guidance is largely left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lyricsA
Get lyrics for a track from LRCLIB. Returns synced lyrics if available.
| Name | Required | Description | Default |
|---|---|---|---|
| track_name | Yes | Track name | |
| artist_name | Yes | Artist name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the behavioral burden. It does disclose the external source and the conditional 'if available' behavior for synced lyrics. However, it omits details about failures, unsynced fallback, rate limits, or any payload structure beyond the synced-lyrics mention.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the core action, the source, and the primary return caveat with no redundant words. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter lookup tool, the description covers the source, what it returns, and the main conditional (synced lyrics if available). Since there is no output schema, stating the return value is important and is handled, though minor gaps like unspecified no-result behavior prevent a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both track_name and artist_name are self-explanatory in the schema. The description adds no extra guidance about parameter format, escaping, or how strict the matching is, so it neither improves nor degrades the schema's clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') with a clear resource ('lyrics for a track') and identifies the data source (LRCLIB). This makes it immediately distinguishable from siblings like search_lyrics, which implies finding rather than exact lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: given a track_name and artist_name, retrieve lyrics. However, the description does not explicitly state when to use this over search_lyrics, nor does it mention conditions or exclusions such as requiring exact artist/track names or preferring search_lyrics when the track is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingA
Check if Synesthesia is running
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It discloses the core behavior (checking if Synesthesia is running) but does not specify the return format or behavior when the service is down (e.g., error vs. false). This is a minor gap for a zero-parameter health check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the core purpose immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter health check with no output schema, the description is nearly complete. The only missing detail is what the response indicates, but an agent can invoke it correctly without that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so the baseline of 4 applies. The description correctly adds no unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and resource ('Synesthesia'), clearly identifying it as a health-check tool. This distinguishes it from the sibling tools that analyze YouTube, download audio, or retrieve lyrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to use this tool versus the sibling tools, nor does it mention that it could be a prerequisite liveness check before other operations. The usage context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_lyricsB
Search for lyrics on LRCLIB
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query (track name, artist, or lyrics) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full responsibility for behavioral disclosure. It does not mention what the search returns, whether partial matches are allowed, or any limits or ordering, providing no behavioral context beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; every word earns its place. The purpose is front-loaded and the description is appropriately sized for a one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the one parameter is well documented, but there is no output schema and no indication of what a search response looks like. The lack of usage guidance also leaves ambiguity about when to call this instead of get_lyrics, making this minimally viable but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully describes the query parameter (100% coverage), so the baseline is 3. The description adds no additional parameter guidance, but the schema's note that the query can be a track name, artist, or lyrics is sufficient for invoking the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search') and resource ('lyrics on LRCLIB'), making the tool's purpose clear. It doesn't explicitly contrast with sibling get_lyrics, but the word 'search' distinguishes it from a direct retrieval tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus get_lyrics, analyze_youtube, or download_audio. The description only states what the tool does, leaving the selection decision entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
5 tool updates
v1.0.0- First observed
analyze_youtube - First observed
download_audio - First observed
get_lyrics - First observed
ping - First observed
search_lyrics
TDQS
analyze_youtube bundles downloading, audio feature extraction, lyrics fetching, and visualization, so its scope overlaps with download_audio and get_lyrics. search_lyrics and get_lyrics are reasonably distinct but rely on the agent infering LRCLIB search-vs-get semantics. Overall, the one-stop analysis tool blurs boundaries.
Four tools follow a clear verb_object pattern: analyze_youtube, download_audio, get_lyrics, search_lyrics. ping is a minor deviation but is a conventional health-check exception. Names are otherwise predictable and readable.
Five tools is a tight, focused set for a YouTube audio and lyrics utility. Each tool has a functional reason to exist, and the count avoids both bloat and thinness.
The core workflow of downloading YouTube audio, analyzing it, and retrieving lyrics is covered without dead ends. Niche operations like analyzing an already-local file or listing past analyses are missing, but agents can work around these gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Audio features + harmonic set-building for tracks by name/ISRC. Spotify audio-features replacement.
Create and track AI music videos and audio-reactive visuals from songs.
Fetch transcripts, subtitles, chapters, metadata and frames from YouTube and 10+ video platforms
Write lyrics in 100+ styles, score them, generate full songs with 4 engines, split stems. OAuth.
1
Related MCP Servers
- -licenseNot gradedqualityNot gradedmaintenanceEnables LLMs to search, download, and extract information from YouTube music videos, converting them to high-quality MP3 files.-
- AlicenseBqualityCmaintenanceDownloads YouTube audio and transcribes it locally using faster-whisper, saving transcripts as Markdown and JSON files for Obsidian and Hermes ingestion.4Apache 2.0
- FlicenseNot gradedqualityBmaintenanceAnalyzes audio files to extract exact, reproducible measurements like loudness, tempo, key, spectral balance, and clipping for LLM-based DAW control.-
- AlicenseCqualityCmaintenanceSyncs YouTube Music liked songs, analyzes them for DJ metadata like BPM and key, and enables creating playlists from previews.10MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/amarisaster/Synesthesia'
If you have feedback or need assistance with the MCP directory API, please join our Discord server