noisy-coding
Officialnoisy-coding
Sprich mit Claude Code, während er arbeitet — sprachgesteuertes Coden im Jarvis-Stil. Es ist deine Stimme, die lärmt, nicht dein Code.
Claude spricht kurze Zusammenfassungen laut aus. Ein ständig aktiver Zuhörer verwandelt deine Sprache in Nachrichten, die Claude mitten in der Aufgabe erhält, ohne ihn zu stoppen — kein Push-to-Send, kein Kopieren-Einfügen von Transkripten. Tritt von der Tastatur weg und steuere deinen Agenten weiter.
Warum es dir gefallen wird
Unterbrechungsfreier Ablauf — sprich, während Claude arbeitet; deine Wörter landen in der laufenden Sitzung, nicht in einem Textfeld.
Freihändige Überprüfungen — Claude liest seine Ergebnisse laut vor; du antwortest aus dem anderen Zimmer.
Live-Dashboard mit "taktischem HUD" — Gesprächsprotokoll mit Wiedergabe/Abruf, Echtzeit-Oszilloskop, Stummschaltflächen, Kosten und Latenzen auf einen Blick.
Pro-Agent-Charakter — Stimme, Geschwindigkeit und Persönlichkeitsregler für jeden Agenten, alles über das Dashboard.
Nichts in Dateien konfigurieren — API-Schlüssel, Geräte, Sprache, Push-to-Talk: alles lebt in der UI und bleibt erhalten.
Redet dir nie dazwischen — immer nur eine Stimme; Sprache, die du verpasst hast, wird als UNHEARD geparkt, und ein Button CATCH UP spielt sie erneut ab.
Spracherkennung und Sprachsynthese laufen über die Grok (xAI) Voice API — in der Praxis extrem günstig (ein kleines einmaliges Budget reicht für Monate täglicher Nutzung).
Related MCP server: Elba MCP Server
Installation in 2 Minuten
Das Backend wird als Docker-Image ohne spezielle Hardware geliefert (noisy/noisy-coding): Der Browser-Tab des Dashboards ist das Mikrofon und der Lautsprecher. Du brauchst Docker und einen Browser — kein Python, kein Git, keine Umgebungsvariablen.
# terminal: marketplace + plugin in one line
claude plugin marketplace add noisy/noisy-coding && claude plugin install noisy-coding@noisy# inside Claude Code (new session):
/noisy-coding:setupDer Setup-Befehl startet das veröffentlichte Image und führt dich durch den ersten Kontakt. Dann schließt du im Browser unter http://127.0.0.1:8765 ab: Füge deinen xAI-API-Schlüssel ein (console.x.ai) und klicke auf das bernsteinfarbene Banner ENABLE TAB AUDIO — dieser eine Klick macht den Tab zu deinem Mikrofon und Lautsprecher. Lass den Tab geöffnet und sprich einfach.
Bleibst du lieber in Claude Code? Gleiche Sache, vier Befehle:
/plugin marketplace add noisy/noisy-coding →
/plugin install noisy-coding@noisy → /reload-plugins →
/noisy-coding:setup.
Andere Setup-Varianten — reines Docker ohne Plugin, native Installation mit Hardware-Mikrofon/Lautsprechern, entfernte Hosts, alle Konfigurationsoptionen — findest du in docs/INSTALL.md.
So funktioniert's
Die gesamte Sprachlogik lebt in einem Listener-Daemon — dem alleinigen Besitzer von Mikrofon, Wiedergabewarteschlange und Lautsprechern. Der MCP-Server ist ein schlanker Vermittler, der speak-Anfragen weiterleitet; die Hooks von Claude Code liefern deine transkribierte Sprache zurück in die Sitzung (siehe docs/hooks.md).
mic (hardware or browser tab via WS :8766)
-> VAD -> Grok STT -> transcript queue -> HTTP :8765
^ polled by Claude Code hooks
speak (MCP, stdio or HTTP :8767) -> POST /speak -> daemon queue
-> Grok TTS -> speakers (hardware or browser tab)Werkzeuge
Tool | Was es tut |
| Stellt |
| Fire-and-forget-Variante: kehrt sofort zurück, spielt im Hintergrund. |
| Schaltet die Stimme dieses Agenten bewusst um (bleibt erhalten, wird im Dashboard angezeigt). |
| Listet die eingebauten Stimmen von Grok auf ( |
Dokumentation
docs/INSTALL.md — reines Docker, native Installation, entfernte Hosts, Umgebungsvariablen, Entwicklungsbefehle
docs/hooks.md — wie Claude dich hört
docs/ports.md — wofür jeder Port da ist
docs/local-development.md — am noisy-coding selbst herumexperimentieren
Lizenz
MIT © Krzysztof Szumny
Available Tools
4 toolsannounceA
Speak a quick spoken update WITHOUT waiting for it to finish.
Fire-and-forget: use this to tell the user what you just did and keep
working ("done with X, moving on") — it returns immediately and plays in
the background, queued behind any current speech. Use speak instead when
you are asking a question or otherwise waiting for the user's reply.
Like speak, it carries only text — voice/speed/language live in the daemon.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral burden. It discloses that the tool returns immediately, plays in the background, queues behind current speech, and carries only text. This is thorough and gives the agent a clear model of runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured, front-loading the core behavior in the first sentence and building on it with usage guidance. Every sentence adds value, with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with an output schema present, this description covers purpose, usage, behavior, and parameter meaning fully. It is complete and leaves no significant gaps for an agent to operate correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'text' parameter has zero schema documentation, but the description compensates by stating the tool 'carries only text' and that voice/speed/language live in the daemon, clarifying that the text parameter is the complete content with no hidden options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Speak a quick spoken update WITHOUT waiting for it to finish.' It directly distinguishes the tool from its sibling 'speak' by emphasizing the fire-and-forget behavior, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is provided: use this for quick updates while continuing work, and 'Use `speak` instead when you are asking a question or otherwise waiting for the user's reply.' This clearly states when to use this tool versus an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
change_voiceA
Deliberately switch this agent's speaking voice from now on.
Updates your character in the listener daemon: the dashboard shows the new voice and every later speak/announce uses it (it also persists across restarts). Use list_voices to see the options. Speak itself carries no voice information — this call is the only way to change how you sound, so use it consciously (e.g. when the user asks for it).
Args: voice_id: Which voice to switch to. speaker: Move a named SPEAKER's voice instead of your own — the personas you address with speak(speaker=...). A voice already held by someone else is refused rather than duplicated, so two speakers never become indistinguishable by ear.
| Name | Required | Description | Default |
|---|---|---|---|
| speaker | No | ||
| voice_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavior. It discloses that the change persists across restarts, affects the listener daemon, applies to future speak/announce calls, and that duplicate voices are refused to prevent indistinguishable speakers. This dramatically exceeds baseline explanation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized in clear sections with the proactive sentence first, followed by practical details and then argument semantics. Every sentence adds relevant information; there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a persistent state-changing tool with a lightness of given structured metadata. The description covers side effects, persistence, usage context, the valid arguments, and the duplicate-owner failure behavior. There is enough to invoke intentionally and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no parameter descriptions (0% coverage), so the description is the only source of meaning. It explains voice_id as the voice to switch to and elaborates on speaker, including its use for named SPEAKER-defined personas and the duplicate-refusal behavior. It could be slightly stronger on voice_id's allowed values, but overall it compensates well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Deliberately switch this agent's speaking voice from now on' and details the state change. It distinguishes itself from the sibling speech tools by explicitly stating that speak carries no voice information and that this is the only way to change how one sounds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this call when the user asks for a voice change, use list_voices to see options, and don't expect speak/announce to carry voice information. This effectively tells the agent when to use this tool versus its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voicesA
List the Grok TTS voices available for the speak tool.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so description carries full burden. 'List' clearly signals a read-only operation and adds the scope 'for the speak tool'. It does not disclose return format or dynamic behavior, but for a simple listing tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence with a clear verb and object. No redundant information, every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with zero parameters and an output schema present. The description fully covers purpose and intended use, making it complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline score of 4 applies. The description correctly adds no unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'List' with resource 'Grok TTS voices' and clarifies they are for the speak tool, distinguishing from sibling tools that speak or change voice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use before speak to discover available voices, and the phrase 'available for the speak tool' gives context. No explicit exclusions or alternatives, but the purpose is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speakA
Speak a short message aloud to the user through their speakers.
Use this to deliver a spoken TL;DR alongside (not instead of) your written answer: 1-3 conversational sentences summarizing the outcome, a finding, or a question. Never read code, file paths, or long explanations aloud.
You send only the text: voice, speed and language belong to the daemon (the user controls them on the dashboard). To deliberately switch your voice, call change_voice.
Concurrent speech is serialized: by default a new call WAITS for the current utterance to finish (queued), and for the user to finish speaking. Set interrupt=True to cut the current utterance off and speak immediately — use it only when your previous words are now stale (e.g. the user corrected you mid-answer).
Args: text: What to say. Plain conversational prose. Mark the key words the listener must catch with markdown bold (like this) — they get vocal emphasis and show bold on the live dashboard. Also supports inline speech tags like [pause] or [laugh] and wrapping tags like text. interrupt: Cut off any utterance currently playing and speak now. speaker: ONLY for subagents. If you are a subagent (Task/Agent tool), pass your role name here (e.g. "researcher") — the dashboard shows the message under that name with its own portrait, and the daemon gives you a stable voice distinct from the main agent's. The main agent must leave this empty.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| speaker | No | ||
| interrupt | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does so thoroughly: it discloses serialized/concurrent speech queuing, interrupt semantics, that only text is sent (voice/language controlled by daemon), and speaker-role constraints for subagents. This is rich behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured: purpose first, then usage guidelines, then behavioral notes, then parameter details. Every section earns its place, though the opening sentence and the second sentence partially overlap in saying it's a short spoken message.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (queueing, interrupt, subagent speaker, formatting tags), no annotations, and an output schema, the description covers all necessary context: when to use, exclusions, behavior, parameter semantics, and related tools. It is a complete standalone guide.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and does. It explains text formatting (markdown bold, [pause], [laugh], <soft> tags), the exact meaning of interrupt, and the speaker parameter's subagent-only usage with dashboard implications.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description immediately states 'Speak a short message aloud to the user through their speakers,' a specific verb+resource pairing. It further scopes usage to 'a spoken TL;DR alongside (not instead of) your written answer' and explicitly contrasts with change_voice, distinguishing it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: deliver 1-3 conversational sentences summarizing outcome/finding/question, never read code/paths/long explanations. It also names the alternative change_voice for switching voices and explains interrupt behavior, giving clear conditions for interrupt=True.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v2.16.0- Changed
change_voice1 field changed- added
Input schema / properties / speakerAdded value: +{ + "default": "", + "title": "Speaker", + "type": "string" +}
4 tool updates
v2.13.4- First observed
announce - First observed
change_voice - First observed
list_voices - First observed
speak
TDQS
The set is mostly distinct: speak delivers a blocking utterance, announce is fire-and-forget, change_voice and list_voices have clearly separate roles. speak and announce both produce speech, so their overlap could cause an agent to misselect when the blocking behavior matters, but the descriptions offer strong guidance.
Naming is simple and readable with all verbs as commands, but it mixes one-word verb names (speak, announce) with verb_noun patterns (change_voice, list_voices). This minor inconsistency is not confusing and the style remains predictable.
Four tools is well-scoped for a voice/speech server: each tool covers a necessary function—speaking, quick updates, voice switching, and voice enumeration. There is no bloat or obvious missing core capability for the stated purpose.
The tool set covers the main lifecycle of spoken interaction: speak, gently tell what you're doing, switch voices persistently, and discover available voices. One possible gap is lack of a way to query the currently active voice, but this is a minor issue that does not block common workflows.
Maintenance
Related MCP Connectors
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Real-time chat for AI agents. Claude Code, Cursor, Cline and Codex join channels over MCP.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
- AxisOAuthdev.useaxis
Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables bidirectional voice interaction for Claude Code using local speech-to-text and text-to-speech models optimized for Apple Silicon. It provides tools to listen to user speech via microphone and speak responses aloud through system speakers.16Apache 2.0
- AlicenseAqualityDmaintenanceManage voice AI agents from Claude Code, Cursor, VS Code, or any MCP-compatible assistant.3153MIT
- AlicenseAqualityDmaintenanceLocal speech-to-text transcription using Microsoft's VibeVoice-ASR model with speaker diarization, enabling audio transcription directly in AI tools like Claude Code, Cursor, and OpenCode.32MIT
- FlicenseNot gradedqualityBmaintenanceProvides bidirectional local voice for Claude Code on Apple Silicon, enabling hands-free conversation and spoken replies using local Whisper STT and Kokoro TTS, with optional ElevenLabs backend and a Stop hook for automatic speech.-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/noisy/noisy-coding'
If you have feedback or need assistance with the MCP directory API, please join our Discord server