embecode
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@embecodesearch for function that handles user authentication"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
embecode
Local-first MCP server for semantic + keyword hybrid code search. Zero external services. No API keys required.
Usage
# From your project root
uvx embecode
# Or with an explicit path
uvx embecode --path /path/to/repoAdd to your MCP client config (Claude Desktop, Cursor, Cline, etc.):
{
"mcpServers": {
"embecode": {
"command": "uvx",
"args": ["embecode"]
}
}
}Related MCP server: reporelay
Tools
Tool | Description |
| Hybrid semantic + keyword search over your codebase |
| Check indexing progress, file count, and last updated time |
How it works
Parses files into AST chunks via tree-sitter (cAST algorithm)
Embeds chunks locally with sentence-transformers (
nomic-embed-text-v1.5)Stores vectors + FTS index in a single DuckDB file at
~/.cache/embecode/Fuses BM25 and dense vector results with Reciprocal Rank Fusion
Watches for file changes via watchfiles and re-indexes incrementally
Development
# Install dependencies
uv sync
# Run tests
uv run pytest
# Lint and format
uv run ruff check src/ tests/
uv run ruff format src/ tests/Benchmarks
Two benchmark classes live in tests/test_performance.py and use pytest-benchmark:
Class | DB | What it measures |
| Mock (in-memory dict) |
|
| Real DuckDB (VSS + FTS) | Actual query latency: cosine-similarity scan, BM25, and fusion |
Run the real benchmarks:
pytest tests/test_performance.py::TestSearchBenchmarkReal -v --benchmark-only --no-cov -sThe first run builds a 200-file synthetic index into .bench_db/ (~20s). Subsequent runs reuse it and start immediately. Delete .bench_db/ to force a rebuild.
Run the mock benchmarks (no setup cost, useful for isolating Searcher logic overhead):
pytest tests/test_performance.py::TestSearchBenchmark -v --benchmark-only --no-cov -sReading the output:
Each test prints a per-phase timing breakdown from SearchTimings on the last benchmark round:
phase breakdown (last run): {'embedding_ms': 0.0, 'vector_search_ms': 78.5, 'bm25_search_ms': 6.5, 'fusion_ms': 0.01, 'total_ms': 85.0}pytest-benchmark then prints a summary table with min, max, mean, median, and stddev across all rounds.
Requires Python 3.12.
Available Tools
2 toolsindex_statusA
Get current index status.
Returns: Dictionary with: - files_indexed: Number of files that have been indexed - total_chunks: Total number of chunks in the index - embedding_model: Name of the embedding model in use - last_updated: ISO timestamp of last index update - is_indexing: Whether indexing is currently in progress - current_file: Current file being indexed (if is_indexing=True) - progress: Progress as a fraction 0-1 (if is_indexing=True)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently lists return fields, including conditional fields like current_file and progress that appear only when is_indexing=True. It also provides semantics (e.g., progress as a fraction 0-1) and notes the timestamp format. It does not explicitly state that it is a read-only operation, but the verb and nature of the tool imply this.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose statement, followed by a bulleted list of return fields. Each bullet adds distinct value, clarifying field names and conditional behavior. There is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and the presence of an output schema, the description is sufficiently complete. It goes beyond the schema by explaining conditional fields and value formats, and it covers all relevant behavioral aspects for a status-checking tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4 per the rubric. The description appropriately does not attempt to explain parameter semantics, as there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Get current index status,' which clearly identifies a specific verb and resource. This distinguishes it from the sibling tool 'search_code' by focusing on index health rather than code search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: call this when you need current index status. However, there is no explicit guidance on when to use this versus alternatives, nor any mention of prerequisites or exclusions, leaving the agent to infer from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codeA
Search the codebase using keyword, semantic, or hybrid search.
Args: query: Search query string (natural language or code). mode: Search mode - "semantic" for vector search, "keyword" for BM25, or "hybrid" for RRF fusion of both (default). top_k: Number of results to return (default: 10). path: Optional path prefix filter (e.g., "src/", "apps/ui/").
Returns: List of concise chunk results with file_path, language, start_line, end_line, definitions, preview, and relevance score.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | hybrid | |
| path | No | ||
| query | Yes | ||
| top_k | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, but the description discloses key behavioral traits: it returns 'concise chunk results' with specific fields (file_path, language, start_line, end_line, definitions, preview, and relevance score), and explains the default mode and return behavior. It does not explicitly state side effects (e.g., read-only), but 'search' inherently implies non-mutation, lending sufficient transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with an Args/Returns format. Every sentence is informative: it states the purpose, enumerates parameters with defaults, and describes the return format. No redundant or filler content; it is appropriately sized for a tool with four parameters and multiple modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description covers all essential aspects: purpose, parameters, defaults, and return value structure (already partially provided via output schema, but description reinforces it). The presence of an output schema and explicit Returns section makes it complete for an agent to invoke correctly. The relationship to index_status is implicit through naming.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description fully compensates by explaining every parameter: query, mode (with enum equivalent wording), top_k (default), and path (with example). It adds semantic meaning beyond the schema, clarifying mode options and path filter semantics, making parameter usage unambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's purpose: 'Search the codebase using keyword, semantic, or hybrid search.' This is a specific verb+resource+modes formulation that immediately distinguishes it from the sibling tool index_status, which is about indexing status, not searching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on how to use the tool, including three search modes with a default, and optional path filtering. It does not explicitly state when not to use this tool or directly name index_status as an alternative, but the sibling relationship is clear. The mode descriptions implicitly guide selection (e.g., hybrid as default).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v0.1.0- First observed
index_status - First observed
search_code
TDQS
The two tools serve entirely distinct purposes: search_code performs queries over the codebase, while index_status reports on the index state. There is no overlap in functionality, making tool selection unambiguous.
Both names use lowercase with underscores and are descriptive, but search_code follows verb_noun structure while index_status is noun_noun. Minor deviation from a strict verb-noun pattern, but still consistent in style.
With only 2 tools, the server feels thin for a code search/indexing domain. However, the scope is narrow and the tools cover the essential query and status check, so the count is borderline but acceptable.
The tools enable searching and checking index status, but lack operations like triggering reindexing, managing index configuration, or retrieving raw file contents. These gaps may require external processes or additional functionality.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
Capability registry for the agentic economy. Semantic search over verified MCP server listings.
Personal knowledge base MCP server with semantic search, auto-categorization, metadata extraction
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP server with local vector search for your codebase. Smart indexing, semantic search, Git history — all offline.748MIT
- AlicenseNot gradedqualityFmaintenanceSelf-hosted MCP server for indexing and searching code repositories via hybrid search and deep code understanding.2118MIT
- AlicenseAqualityAmaintenanceLocal MCP server for semantic code search using Tree-sitter AST parsing, local embeddings, and hybrid search; enables indexing and querying codebases entirely offline.5MIT
- FlicenseAqualityBmaintenanceSelf-hosted hybrid code search MCP server with text, symbol, and semantic search layers. Runs locally, no third-party MCP servers, LSP, or SaaS.8-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/jdtzmn/embecode'
If you have feedback or need assistance with the MCP directory API, please join our Discord server