Skip to main content
Glama
djelia-org

Djelia MCP Server

Official
by djelia-org

๐ŸŽ™๏ธ Djelia MCP Server

An MCP server for Djelia โ€” bring Bambara transcription, translation, and text-to-speech to any LLM.

Built with FastMCP v3 ยท Python 3.11+ ยท uv-managed


โœจ Overview

Djelia is a linguistic-AI platform focused on African languages โ€” currently Bambara (bam_Latn), with translation bridging to French (fra_Latn) and English (eng_Latn).

This server wraps the Djelia REST API behind the Model Context Protocol, so any MCP-compatible client (Claude Desktop, Cursor, Cline, your own agent) can call Djelia's models as native tools โ€” no SDK glue, no HTTP plumbing in your prompt.

What you get

#

Tool

Direction

V1 / V2

Returns

1

list_supported_languages

โ€”

โ€”

JSON list

2

translate

text โ†’ text

v1

translated text

3

transcribe

audio โ†’ text

v2

text + segment timing

4

text_to_speech

text โ†’ audio

v2

audio content block

Design note: V2 APIs are exposed for transcription and TTS because they supersede V1 (richer voices via description, format control). True /stream endpoints are omitted โ€” MCP is request/response, so we aggregate the stream inside the tool. Add raw streaming tools only if a use case needs them.


Related MCP server: sarvam-tools

๐Ÿ—๏ธ Architecture

flowchart LR
    subgraph Client["MCP Client"]
        LLM["LLM / Agent<br/>(Claude, Cursor, โ€ฆ)"]
    end

    subgraph Server["djelia-mcp-server (this repo)"]
        MCP["FastMCP Server<br/><i>4 tools, stdio ยท sse ยท http</i>"]
        HANDLERS["Tool Handlers<br/>translate ยท transcribe ยท tts"]
        HTTP["httpx.AsyncClient<br/><i>x-api-key header</i>"]
        MCP --> HANDLERS --> HTTP
    end

    subgraph Djelia["Djelia Cloud API"]
        T1["/v1/translate"]
        T2["/v2/transcribe"]
        T3["/v2/tts"]
    end

    LLM -- "MCP JSON-RPC" --> MCP
    HTTP -- "HTTPS" --> T1
    HTTP -- "HTTPS" --> T2
    HTTP -- "HTTPS" --> T3

Key design choices

  • One shared HTTP client โ€” x-api-key header injected once per request; key read from DJELIA_API_KEY env var.

  • base64 for audio input โ€” MCP payloads are JSON; audio bytes travel as base64 so it works across any client. A magic-byte sniffer (_guess_ext) recovers the right file extension for the multipart upload.

  • Audio output as a content block โ€” FastMCP's Audio helper returns a proper MCP audio block (clients receive it base64-encoded).


๐Ÿ”ง How each tool works

1 ยท list_supported_languages

Returns the language codes you'll pass to translate.

sequenceDiagram
    participant C as Client
    participant S as MCP Server
    participant D as Djelia API
    C->>S: list_supported_languages()
    S->>D: GET /api/v1/models/translate/supported-languages
    D-->>S: [{code, name}, ...]
    S-->>C: structured list

2 ยท translate

sequenceDiagram
    participant C as Client
    participant S as MCP Server
    participant D as Djelia API
    C->>S: translate(source, target, text)
    S->>D: POST /api/v1/models/translate (JSON)
    D-->>S: { "text": "<translated>" }
    S-->>C: structured dict

Parameters

Name

Type

Values

source

enum

bam_Latn ยท fra_Latn ยท eng_Latn

target

enum

bam_Latn ยท fra_Latn ยท eng_Latn

text

string

the text to translate

3 ยท transcribe (Bambara audio โ†’ text)

The tool decodes base64 โ†’ sniffs the format โ†’ uploads as multipart to the V2 transcription endpoint.

sequenceDiagram
    participant C as Client
    participant S as MCP Server
    participant D as Djelia API
    C->>S: transcribe(audio_base64)
    S->>S: base64decode + guess_ext (mp3/wav/m4a/ogg)
    S->>D: POST /api/v2/models/transcribe (multipart)
    alt single text response
        D-->>S: { "text": "..." }
    else segmented response
        D-->>S: [{ text, start, end }, ...]
    end
    S-->>C: ToolResult (structured + text)

4 ยท text_to_speech (text โ†’ Bambara audio)

sequenceDiagram
    participant C as Client
    participant S as MCP Server
    participant D as Djelia API
    C->>S: text_to_speech(text, description, format)
    S->>D: POST /api/v2/models/tts (JSON)
    D-->>S: binary audio bytes
    S-->>C: Audio content block (base64)

Parameters

Name

Type

Values

text

string

text to synthesize

description

string

voice style, e.g. "calm male voice, slow pace"

format

enum

mp3 (default) ยท wav ยท wav_8k ยท ulaw_8k


๐Ÿš€ Quickstart

1 ยท Prerequisites

2 ยท Install dependencies

git clone <your-repo-url> djelia-mcp-server
cd djelia-mcp-server
uv sync

3 ยท Set your API key

cp .env.example .env
# edit .env:
#   DJELIA_API_KEY=your_key_here

The server reads DJELIA_API_KEY from the environment. It fails fast with a clear message if the key is missing.


๐ŸŒ Transports

FastMCP supports three transports. Pick the one your client expects.

flowchart TB
    subgraph "Transport decision"
        STDIO["stdio<br/><b>default</b><br/>Claude Desktop, CLI agents"]
        SSE["sse<br/><b>legacy</b><br/>older MCP clients"]
        HTTP["http / streamable-http<br/><b>recommended for network</b>"]
    end
    STDIO -. "stdin/stdout" .-> Srv["FastMCP Server"]
    SSE   -. "HTTP + EventSource<br/>GET /sse/" .-> Srv
    HTTP  -. "HTTP POST<br/>POST /mcp/" .-> Srv

Mode

Command

Endpoint

stdio (default)

uv run fastmcp run server.py

โ€”

sse (legacy)

uv run fastmcp run server.py -t sse -p 8000

http://127.0.0.1:8000/sse/

http

uv run fastmcp run server.py -t http -p 8000

http://127.0.0.1:8000/mcp/

streamable-http

uv run fastmcp run server.py -t streamable-http -p 8000

http://127.0.0.1:8000/mcp/

Override host/port with --host / -p. See all options: uv run fastmcp run --help.

Direct Python (without the fastmcp CLI)

Transport is read from DJELIA_TRANSPORT (stdio | sse | http):

DJELIA_TRANSPORT=sse DJELIA_HOST=127.0.0.1 DJELIA_PORT=8000 uv run python server.py

๐Ÿค Client configuration

Claude Desktop / Cursor (stdio)

Drop this into your MCP client config:

{
  "mcpServers": {
    "djelia": {
      "command": "uv",
      "args": [
        "run",
        "--directory",
        "/absolute/path/to/djelia-mcp-server",
        "fastmcp",
        "run",
        "server.py"
      ],
      "env": {
        "DJELIA_API_KEY": "your_api_key"
      }
    }
  }
}

Remote / networked client (SSE or HTTP)

Run the server with -t sse or -t http, then point your client at the endpoint (e.g. http://your-host:8000/mcp/).


๐Ÿ—‚๏ธ Project layout

djelia-mcp-server/
โ”œโ”€โ”€ server.py         # all 4 tools + httpx client + transport switch
โ”œโ”€โ”€ pyproject.toml    # uv project (fastmcp + httpx)
โ”œโ”€โ”€ .env.example      # DJELIA_API_KEY template
โ”œโ”€โ”€ .gitignore
โ””โ”€โ”€ README.md

One file of code โ€” by design. Tools are co-located because they share one client and one concern (calling Djelia).


๐Ÿงช Verifying it works

Smoke-test that all tools register and the server boots on every transport:

# list registered tools
uv run python -c "import asyncio, server; \
  [print(' -', t.name) for t in asyncio.run(server.mcp.list_tools())]"

# boot a transport
uv run fastmcp run server.py -t sse -p 8000

You should see 4 tools listed, and the FastMCP banner with transport 'sse' followed by Uvicorn running.


๐Ÿ“š References


๐Ÿ“ License

MIT

Available Tools

4 tools
list_supported_languagesA

List languages supported by Djelia translation.

Returns a list of {"code": str, "name": str}. Codes: bam_Latn (Bambara), fra_Latn (French), eng_Latn (English).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. It discloses the return shape (list of {'code','name'}) and provides exact codes and names, which is concrete behavioral information. It omits error handling and auth, but that is acceptable for a simple read-only listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences: first states the purpose plainly, second provides return format and examples. No wasted words; fully front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description fully covers the purpose and return shape. Nothing is missing for an agent to call it correctly; the examples make the output concrete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so baseline is 4. The description adds no parameter info because none is needed; the schema already shows no properties.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses specific verb 'List' and resource 'languages supported by Djelia translation', distinguishing it from siblings (translate, transcribe, text-to-speech) as a discovery tool. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context: use this to see which language codes/names are available for translation. It does not explicitly state alternatives or when-not-to-use, but sibling tools are obviously different, so usage intent is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speechA

Synthesize Bambara speech from text with desired voice description.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to convert to speech.
formatNoOutput audio format. Default mp3.mp3
descriptionYesVoice style/characteristics (e.g. "calm male voice, slow pace").

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only restates the core action without detailing output format, side effects, authentication needs, rate limits, or any constraints. The schema provides the `format` parameter, but the description does not explain what the tool returns or how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that immediately conveys the tool's purpose and key inputs. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool and the well-documented schema, the description is adequate but not complete. It does not mention the return value (e.g., audio data) or any caveats, and the absence of an output schema means more responsibility falls on the description to explain behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds little beyond naming `text` and `description`; it does not clarify the `format` parameter or provide any additional meaning beyond the existing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Synthesize') and resource ('Bambara speech'), and clarifies the two main inputs (`text` and `description`). This clearly distinguishes it from sibling tools like `translate`, `transcribe`, and `list_supported_languages`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for converting text to Bambara speech, but it does not explicitly state when to use it over alternatives or mention any exclusions. There is no guidance on when not to use it or when a sibling tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribeB

Transcribe Bambara speech to text using Djelia V2.

ParametersJSON Schema
NameRequiredDescriptionDefault
audio_base64YesAudio file bytes encoded as base64. Supported: common formats (mp3, wav, m4a, ...).

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only mentions the model name and task, without detailing output format, supported audio formats beyond what the schema implies, or any limitations. This is minimal behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the action and subject. It contains no filler, every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is expected to provide more context about return values, error conditions, or behavioral expectations. It offers none, leaving significant gaps for an agent trying to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides a complete description of the single parameter (audio_base64) with format details, achieving 100% schema coverage. The tool description adds no additional parameter-specific meaning, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'transcribe' and clearly identifies the resource ('Bambara speech') and the model ('Djelia V2'), distinguishing it from sibling tools like translate and text_to_speech.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for transcribing Bambara audio, giving clear context, but it does not explicitly state when to use it versus alternatives or mention any exclusions. Sibling tools are provided, but the description itself lacks direct guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

translateA

Translate text from source to target language.

Use list_supported_languages to get valid codes. Returns {"text": ""}.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
sourceYes
targetYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the return format as a JSON object with the translated text, which is useful behavioral context. There are no annotations, but it does not cover error behavior, restrictions, or whether source and target must differ, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, consisting of two sentences. It fronts the main action and includes only essential additional guidance about language codes and return format, with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple translation tool with three parameters and no annotations, the description covers the core purpose, parameter roles, and return structure. It does not discuss edge cases like invalid codes or source-target equality, but the pointer to list_supported_languages helps mitigate the need for more detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explains the role of each parameter (text, source, target) in context, but schema description coverage is 0%. It does not explain the meanings of the enum values (e.g., bam_Latn), relying on a pointer to list_supported_languages, which partially compensates for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Translate') and operands (text, source, target language). It distinguishes itself from sibling tools such as transcribe and text_to_speech by specifying the exact task of language translation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (translation) and explicitly directs users to list_supported_languages for valid codes, establishing a prerequisite. However, it does not mention when to avoid this tool or explicitly compare with alternatives like transcribe.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv0.1.0
    • First observedlist_supported_languages
    • First observedtext_to_speech
    • First observedtranscribe
    • First observedtranslate

TDQS

A3.9/5.0
Disambiguation5/5

Each tool targets a distinct operation: listing languages, translating text, transcribing speech, and synthesizing speech. There is no overlap or ambiguity between them.

Naming Consistency3/5

Tool names mix verb-led formats like 'list_supported_languages' and 'translate' with the noun phrase 'text_to_speech'. The pattern is not fully consistent but remains readable and predictable.

Tool Count5/5

Four tools is a well-scoped size for a language services server, covering translation, transcription, TTS, and language discovery without bloat or thinness.

Completeness4/5

The tool surface covers the core language lifecycle: discover languages, translate, transcribe, and synthesize. Minor gaps exist (e.g., no explicit voice listing or language detection), but they are not obvious dead ends for the stated purpose.

Maintenance

ActivityStale
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/djelia-org/djelia-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server