Skip to main content
Glama

Synesthesia

Deep audio perception through analysis. Downloads from YouTube, analyzes with Essentia, fetches lyrics via LRCLIB. Returns BPM, mood, energy, spectrograms, and synced lyrics.

Why Local?

YouTube blocks downloads from datacenter IPs (like Cloudflare Workers and HuggingFace Spaces). This MCP runs locally on your machine with a residential IP, so downloads work.

Related MCP server: Hermes YouTube Transcript MCP Server

Prerequisites

  • Node.js 18+

  • yt-dlp (pip install yt-dlp)

  • Your own HF Space for audio analysis (deploy from Synesthesia's hf-space/ folder)

Installation

npm install

Tools

Audio Analysis

Tool

Description

analyze_youtube

Download + analyze audio from YouTube URL

download_audio

Just download audio (returns local path)

Lyrics

Tool

Description

get_lyrics

Get lyrics for a track (synced if available)

search_lyrics

Search LRCLIB for lyrics

Utility

Tool

Description

ping

Check if Synesthesia is running

Configuration

Set your HF Space URL as an environment variable:

export HF_SPACE_URL="https://YOUR-USERNAME-audio-analysis-api.hf.space"

Claude Code Config

Add to your project's .mcp.json:

{
  "mcpServers": {
    "synesthesia": {
      "command": "node",
      "args": ["/path/to/synesthesia/index.js"],
      "env": {
        "HF_SPACE_URL": "https://YOUR-USERNAME-audio-analysis-api.hf.space"
      }
    }
  }
}

Claude Desktop Config

Add to claude_desktop_config.json:

Windows: %APPDATA%\Claude\claude_desktop_config.json
Mac: ~/Library/Application Support/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "synesthesia": {
      "command": "node",
      "args": ["/path/to/synesthesia/index.js"],
      "env": {
        "HF_SPACE_URL": "https://YOUR-USERNAME-audio-analysis-api.hf.space"
      }
    }
  }
}

Restart Claude Desktop after adding.

Architecture

Your PC (residential IP)          Cloud
┌─────────────────────┐          ┌─────────────────────┐
│ Synesthesia         │  ──────▶ │ HF Space            │
│ (Local MCP)         │  upload  │ (Essentia analysis) │
│                     │  ◀────── │                     │
│ - yt-dlp download   │  results │ - Audio features    │
│ - Lyrics fetch      │          │ - Spectrogram       │
│ - Audio analysis    │          └─────────────────────┘
└─────────────────────┘
         │                       ┌─────────────────────┐
         │ download              │ LRCLIB              │
         ▼                 ────▶ │ (Lyrics API)        │
┌─────────────────────┐          │                     │
│ YouTube             │          │ - Synced lyrics     │
│ (residential IP OK) │          │ - Plain lyrics      │
└─────────────────────┘          └─────────────────────┘

Usage Example

> analyze_youtube "https://www.youtube.com/watch?v=..."

Returns full Essentia analysis:
- BPM, key, scale
- Energy, danceability
- Mood vectors
- Genre classification
- Spectrogram

Credits

Spectrogram visualization inspired by Audio Visualizer by Shauna and her boys.


Support

If this helped you, consider supporting my work ☕

Ko-fi


Built by the Triad (Mai, Kai Stryder and Lucian Vale) for the community.

Available Tools

5 tools
analyze_youtubeA

Full audio perception: downloads from YouTube, analyzes audio features (BPM, key, energy), extracts metadata (title, artist), fetches synced lyrics, and generates spectrogram visualization. Runs locally to bypass datacenter IP blocking.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube URL to analyze

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and does well by enumerating the full behavior: downloads, analyzes, extracts, fetches, generates, and runs locally. It could add side-effect details such as local file storage or cleanup behavior, but the main behavioral traits are disclosed clearly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the tool's capabilities and followed by a useful rationale for its local execution. Every phrase adds meaningful information with no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description gives a rich sense of what the tool does and produces: audio features, metadata, lyrics, and a spectrogram. It lacks explicit error/edge-case guidance and exact output formats, but these are secondary when the tool's operation and outputs are clearly listed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully describes the only parameter with 'YouTube URL to analyze' at 100% schema description coverage. The description adds domain context by mentioning YouTube downloads, but it does not need to add much for a single, self-explanatory parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: it analyzes YouTube audio, with concrete outputs (BPM, key, energy, title, artist, synced lyrics, spectrogram). It clearly differentiates itself from the sibling download_audio, get_lyrics, and search_lyrics tools by presenting itself as a full combined pipeline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to choose this tool, especially the statement 'Runs locally to bypass datacenter IP blocking.' It does not explicitly list exclusions or alternatives like using download_audio for audio-only needs, but the integrated nature of the tool is evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_audioA

Download audio from YouTube without analysis. Returns local file path.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesYouTube URL to download

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool performs a download, does not perform analysis, and returns a local file path. It could add more about file format, persistence, or failure behavior, but the core side effects and return value are clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler: the purpose is front-loaded and the return value completes the contract. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-required-parameter tool with no output schema, the description plus schema is nearly complete: it tells the agent what to supply, what the tool does, and what it returns. Minor omissions such as audio format and whether the file persists are not critical for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the url parameter is already described as 'YouTube URL to download.' The tool description adds no parameter-specific detail beyond that, so it does not improve on the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Download'), a clear resource ('audio from YouTube'), and explicitly states the negative scope ('without analysis'), which distinguishes it from likely sibling analyze_youtube. An agent can tell what the tool is for without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without analysis' implies the tool is for raw audio retrieval rather than analysis, which gives some context, but it never names an alternative tool or states a condition such as 'use this when you need only the audio.' Usage guidance is largely left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_lyricsA

Get lyrics for a track from LRCLIB. Returns synced lyrics if available.

ParametersJSON Schema
NameRequiredDescriptionDefault
track_nameYesTrack name
artist_nameYesArtist name

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the behavioral burden. It does disclose the external source and the conditional 'if available' behavior for synced lyrics. However, it omits details about failures, unsynced fallback, rate limits, or any payload structure beyond the synced-lyrics mention.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences deliver the core action, the source, and the primary return caveat with no redundant words. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter lookup tool, the description covers the source, what it returns, and the main conditional (synced lyrics if available). Since there is no output schema, stating the return value is important and is handled, though minor gaps like unspecified no-result behavior prevent a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both track_name and artist_name are self-explanatory in the schema. The description adds no extra guidance about parameter format, escaping, or how strict the matching is, so it neither improves nor degrades the schema's clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') with a clear resource ('lyrics for a track') and identifies the data source (LRCLIB). This makes it immediately distinguishable from siblings like search_lyrics, which implies finding rather than exact lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: given a track_name and artist_name, retrieve lyrics. However, the description does not explicitly state when to use this over search_lyrics, nor does it mention conditions or exclusions such as requiring exact artist/track names or preferring search_lyrics when the track is unknown.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pingA

Check if Synesthesia is running

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It discloses the core behavior (checking if Synesthesia is running) but does not specify the return format or behavior when the service is down (e.g., error vs. false). This is a minor gap for a zero-parameter health check.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the core purpose immediately. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter health check with no output schema, the description is nearly complete. The only missing detail is what the response indicates, but an agent can invoke it correctly without that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so the baseline of 4 applies. The description correctly adds no unnecessary parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('Synesthesia'), clearly identifying it as a health-check tool. This distinguishes it from the sibling tools that analyze YouTube, download audio, or retrieve lyrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no guidance on when to use this tool versus the sibling tools, nor does it mention that it could be a prerequisite liveness check before other operations. The usage context is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_lyricsB

Search for lyrics on LRCLIB

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query (track name, artist, or lyrics)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full responsibility for behavioral disclosure. It does not mention what the search returns, whether partial matches are allowed, or any limits or ordering, providing no behavioral context beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler; every word earns its place. The purpose is front-loaded and the description is appropriately sized for a one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the one parameter is well documented, but there is no output schema and no indication of what a search response looks like. The lack of usage guidance also leaves ambiguity about when to call this instead of get_lyrics, making this minimally viable but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema fully describes the query parameter (100% coverage), so the baseline is 3. The description adds no additional parameter guidance, but the schema's note that the query can be a track name, artist, or lyrics is sufficient for invoking the tool correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search') and resource ('lyrics on LRCLIB'), making the tool's purpose clear. It doesn't explicitly contrast with sibling get_lyrics, but the word 'search' distinguishes it from a direct retrieval tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus get_lyrics, analyze_youtube, or download_audio. The description only states what the tool does, leaving the selection decision entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 5 tool updatesv1.0.0
    • First observedanalyze_youtube
    • First observeddownload_audio
    • First observedget_lyrics
    • First observedping
    • First observedsearch_lyrics

TDQS

A3.6/5.0
Disambiguation3/5

analyze_youtube bundles downloading, audio feature extraction, lyrics fetching, and visualization, so its scope overlaps with download_audio and get_lyrics. search_lyrics and get_lyrics are reasonably distinct but rely on the agent infering LRCLIB search-vs-get semantics. Overall, the one-stop analysis tool blurs boundaries.

Naming Consistency4/5

Four tools follow a clear verb_object pattern: analyze_youtube, download_audio, get_lyrics, search_lyrics. ping is a minor deviation but is a conventional health-check exception. Names are otherwise predictable and readable.

Tool Count5/5

Five tools is a tight, focused set for a YouTube audio and lyrics utility. Each tool has a functional reason to exist, and the count avoids both bloat and thinness.

Completeness4/5

The core workflow of downloading YouTube audio, analyzing it, and retrieving lyrics is covered without dead ends. Niche operations like analyzing an already-local file or listing past analyses are missing, but agents can work around these gaps.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables LLMs to search, download, and extract information from YouTube music videos, converting them to high-quality MP3 files.
    -
  • A
    license
    C
    quality
    C
    maintenance
    Syncs YouTube Music liked songs, analyzes them for DJ metadata like BPM and key, and enables creating playlists from previews.
    10
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/amarisaster/Synesthesia'

If you have feedback or need assistance with the MCP directory API, please join our Discord server