Skip to main content
Glama

embecode

Local-first MCP server for semantic + keyword hybrid code search. Zero external services. No API keys required.

CI PyPI Python

Usage

# From your project root
uvx embecode

# Or with an explicit path
uvx embecode --path /path/to/repo

Add to your MCP client config (Claude Desktop, Cursor, Cline, etc.):

{
  "mcpServers": {
    "embecode": {
      "command": "uvx",
      "args": ["embecode"]
    }
  }
}

Related MCP server: reporelay

Tools

Tool

Description

search_code

Hybrid semantic + keyword search over your codebase

index_status

Check indexing progress, file count, and last updated time

How it works

  • Parses files into AST chunks via tree-sitter (cAST algorithm)

  • Embeds chunks locally with sentence-transformers (nomic-embed-text-v1.5)

  • Stores vectors + FTS index in a single DuckDB file at ~/.cache/embecode/

  • Fuses BM25 and dense vector results with Reciprocal Rank Fusion

  • Watches for file changes via watchfiles and re-indexes incrementally

Development

# Install dependencies
uv sync

# Run tests
uv run pytest

# Lint and format
uv run ruff check src/ tests/
uv run ruff format src/ tests/

Benchmarks

Two benchmark classes live in tests/test_performance.py and use pytest-benchmark:

Class

DB

What it measures

TestSearchBenchmark

Mock (in-memory dict)

Searcher + RRF code path only — no real DB or model

TestSearchBenchmarkReal

Real DuckDB (VSS + FTS)

Actual query latency: cosine-similarity scan, BM25, and fusion

Run the real benchmarks:

pytest tests/test_performance.py::TestSearchBenchmarkReal -v --benchmark-only --no-cov -s

The first run builds a 200-file synthetic index into .bench_db/ (~20s). Subsequent runs reuse it and start immediately. Delete .bench_db/ to force a rebuild.

Run the mock benchmarks (no setup cost, useful for isolating Searcher logic overhead):

pytest tests/test_performance.py::TestSearchBenchmark -v --benchmark-only --no-cov -s

Reading the output:

Each test prints a per-phase timing breakdown from SearchTimings on the last benchmark round:

phase breakdown (last run): {'embedding_ms': 0.0, 'vector_search_ms': 78.5, 'bm25_search_ms': 6.5, 'fusion_ms': 0.01, 'total_ms': 85.0}

pytest-benchmark then prints a summary table with min, max, mean, median, and stddev across all rounds.

Requires Python 3.12.

Available Tools

2 tools
index_statusA

Get current index status.

Returns: Dictionary with: - files_indexed: Number of files that have been indexed - total_chunks: Total number of chunks in the index - embedding_model: Name of the embedding model in use - last_updated: ISO timestamp of last index update - is_indexing: Whether indexing is currently in progress - current_file: Current file being indexed (if is_indexing=True) - progress: Progress as a fraction 0-1 (if is_indexing=True)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It transparently lists return fields, including conditional fields like current_file and progress that appear only when is_indexing=True. It also provides semantics (e.g., progress as a fraction 0-1) and notes the timestamp format. It does not explicitly state that it is a read-only operation, but the verb and nature of the tool imply this.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose statement, followed by a bulleted list of return fields. Each bullet adds distinct value, clarifying field names and conditional behavior. There is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and the presence of an output schema, the description is sufficiently complete. It goes beyond the schema by explaining conditional fields and value formats, and it covers all relevant behavioral aspects for a status-checking tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4 per the rubric. The description appropriately does not attempt to explain parameter semantics, as there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with 'Get current index status,' which clearly identifies a specific verb and resource. This distinguishes it from the sibling tool 'search_code' by focusing on index health rather than code search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this when you need current index status. However, there is no explicit guidance on when to use this versus alternatives, nor any mention of prerequisites or exclusions, leaving the agent to infer from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeA

Search the codebase using keyword, semantic, or hybrid search.

Args: query: Search query string (natural language or code). mode: Search mode - "semantic" for vector search, "keyword" for BM25, or "hybrid" for RRF fusion of both (default). top_k: Number of results to return (default: 10). path: Optional path prefix filter (e.g., "src/", "apps/ui/").

Returns: List of concise chunk results with file_path, language, start_line, end_line, definitions, preview, and relevance score.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNohybrid
pathNo
queryYes
top_kNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, but the description discloses key behavioral traits: it returns 'concise chunk results' with specific fields (file_path, language, start_line, end_line, definitions, preview, and relevance score), and explains the default mode and return behavior. It does not explicitly state side effects (e.g., read-only), but 'search' inherently implies non-mutation, lending sufficient transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with an Args/Returns format. Every sentence is informative: it states the purpose, enumerates parameters with defaults, and describes the return format. No redundant or filler content; it is appropriately sized for a tool with four parameters and multiple modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the description covers all essential aspects: purpose, parameters, defaults, and return value structure (already partially provided via output schema, but description reinforces it). The presence of an output schema and explicit Returns section makes it complete for an agent to invoke correctly. The relationship to index_status is implicit through naming.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description fully compensates by explaining every parameter: query, mode (with enum equivalent wording), top_k (default), and path (with example). It adds semantic meaning beyond the schema, clarifying mode options and path filter semantics, making parameter usage unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: 'Search the codebase using keyword, semantic, or hybrid search.' This is a specific verb+resource+modes formulation that immediately distinguishes it from the sibling tool index_status, which is about indexing status, not searching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on how to use the tool, including three search modes with a default, and optional path filtering. It does not explicitly state when not to use this tool or directly name index_status as an alternative, but the sibling relationship is clear. The mode descriptions implicitly guide selection (e.g., hybrid as default).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedindex_status
    • First observedsearch_code

TDQS

A4.2/5.0
Disambiguation5/5

The two tools serve entirely distinct purposes: search_code performs queries over the codebase, while index_status reports on the index state. There is no overlap in functionality, making tool selection unambiguous.

Naming Consistency4/5

Both names use lowercase with underscores and are descriptive, but search_code follows verb_noun structure while index_status is noun_noun. Minor deviation from a strict verb-noun pattern, but still consistent in style.

Tool Count3/5

With only 2 tools, the server feels thin for a code search/indexing domain. However, the scope is narrow and the tools cover the essential query and status check, so the count is borderline but acceptable.

Completeness3/5

The tools enable searching and checking index status, but lack operations like triggering reindexing, managing index configuration, or retrieving raw file contents. These gaps may require external processes or additional functionality.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    Self-hosted MCP server for indexing and searching code repositories via hybrid search and deep code understanding.
    21
    18
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Local MCP server for semantic code search using Tree-sitter AST parsing, local embeddings, and hybrid search; enables indexing and querying codebases entirely offline.
    5
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jdtzmn/embecode'

If you have feedback or need assistance with the MCP directory API, please join our Discord server