MCP RAG Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP RAG Serversearch my local documentation for the API authentication flow"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP RAG Server
A Model Context Protocol (MCP) server that provides RAG (Retrieval-Augmented Generation) functionality using local embeddings via Ollama and Chroma vector database.
Features
Local Processing: No external API costs - runs entirely locally
Multiple Formats: Supports PDF, Markdown, and TXT files
Smart Chunking: Configurable chunk size with overlap for better context
Vector Search: Semantic search using nomic-embed-text model via Ollama
MCP Integration: Works seamlessly with Cursor and other MCP clients
Related MCP server: MCP RAG with ChromaDB
Prerequisites
Node.js (v18 or higher)
Docker (for ChromaDB)
Homebrew (for Ollama on macOS)
🚀 Quick Start
Setup (one time)
npm run setupThis will:
Start Ollama and install nomic-embed-text model
Start ChromaDB with Docker
Build the project
Ingest documents from
./docs
Development
# Start MCP server
npm run dev
# Ingest new documents
npm run ingestStop Services
npm run stopConfiguration
The server uses a config.json file for configuration:
{
"documentsPath": "./docs",
"chunkSize": 1000,
"chunkOverlap": 200,
"ollamaUrl": "http://localhost:11434",
"embeddingModel": "nomic-embed-text",
"chromaUrl": "http://localhost:8001",
"collectionName": "rag_documents",
"mcpServer": {
"name": "mcp-rag-server",
"version": "1.0.0"
}
}MCP Tools
ingest_docs({path?})- Ingest documents from a directorysearch({query, k?})- Search for relevant document chunksget_chunk({id})- Retrieve a specific chunk by IDrefresh_index()- Clear and refresh the entire index
MCP Resources
rag://collection/summary- Collection statistics and metadatarag://doc/<filename>#<chunk_id>- Individual document chunks
Configure in Cursor
Add to your Cursor MCP settings:
{
"mcpServers": {
"rag-server": {
"command": "node",
"args": ["/Users/luizsoares/Documents/buildaz/mcp_rag/dist/index.js"],
"env": {}
}
}
}Available Scripts
npm run setup- Complete setup (Ollama + ChromaDB + build + ingest)npm run dev- Start MCP server in development modenpm run ingest- Ingest documentsnpm run build- Build the projectnpm run test- Run testsnpm run stop- Stop all services
Troubleshooting
Ollama Connection Issues: Ensure Ollama is running on the configured URL
Model Not Found: Run
ollama pull nomic-embed-textto install the embedding modelDocker Issues: Ensure Docker is running and accessible
Available Tools
4 toolsget_chunkC
Retrieve a specific document chunk by its ID
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The unique identifier of the chunk to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves a chunk, implying a read-only operation, but doesn't mention error handling (e.g., what happens if the ID is invalid), performance characteristics (e.g., speed, caching), or authentication needs. For a retrieval tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that directly states the tool's purpose without any fluff or redundancy. It's appropriately sized for a simple retrieval tool and front-loaded with the essential information, making it highly efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema, no annotations), the description is minimal but adequate for basic understanding. However, it lacks context about what a 'document chunk' is (e.g., part of a larger document), how IDs are obtained, or what the return format looks like (since no output schema exists). For a retrieval tool, this leaves the agent with incomplete information to use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'id' parameter fully documented as 'The unique identifier of the chunk to retrieve'. The description adds no additional meaning beyond this, as it only restates that retrieval is by ID. With high schema coverage, the baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't detract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('retrieve') and resource ('document chunk'), specifying it's by ID. It distinguishes from 'search' (which likely finds chunks by content) and 'ingest_docs' (which adds documents), but doesn't explicitly differentiate from 'refresh_index' (which might update metadata). The purpose is specific but could be slightly more precise about what a 'chunk' entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'search' (for finding chunks by query) or 'refresh_index' (for updating indices). The description implies usage when you have a specific chunk ID, but it doesn't state prerequisites (e.g., needing to know the ID from prior operations) or exclusions (e.g., not for bulk retrieval).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_docsA
Re-ingest documents from the configured documents directory. Use this if search returns no results or if documents have been updated
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the action ('re-ingest') but doesn't disclose behavioral traits like whether this is a read-only or destructive operation, what permissions are required, or how long it might take. The description adds some context about triggers but lacks depth on operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with two sentences: the first states the purpose, and the second provides usage guidelines. Every sentence adds value without unnecessary details, making it front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (no parameters, no output schema, no annotations), the description is somewhat complete but could be improved. It explains when to use the tool but doesn't cover what happens during re-ingestion (e.g., overwrites, indexing behavior) or potential side effects, leaving gaps in understanding for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add parameter semantics, but this is appropriate given the lack of parameters, warranting a baseline score of 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Re-ingest documents from the configured documents directory.' This specifies the verb ('re-ingest') and resource ('documents'), though it doesn't explicitly differentiate from sibling tools like 'refresh_index' which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: 'Use this if search returns no results or if documents have been updated.' This clearly states when to use the tool, offering practical scenarios that help distinguish it from alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_indexB
Clear and refresh the entire document index
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It states the tool 'Clear[s] and refresh[es]' which implies a destructive/replacement operation, but doesn't disclose important behavioral traits like whether this requires admin permissions, how long it takes, if it's asynchronous, what happens during the process, or what the expected outcome is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core purpose without any wasted words. It's appropriately sized for a zero-parameter tool and front-loads the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a potentially destructive operation with no annotations and no output schema, the description is inadequate. It doesn't explain what 'clear' means (deletion? reset?), what 'refresh' entails, whether this affects search availability during the process, or what confirmation/status is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the baseline is 4. The description appropriately doesn't waste space discussing non-existent parameters, though it could theoretically mention that no configuration options are available for this operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Clear and refresh') and target ('the entire document index'), providing a specific verb+resource combination. However, it doesn't differentiate from sibling tools like 'ingest_docs' or 'search' that might also interact with the document index in different ways.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like 'ingest_docs' (which might add documents) or 'search' (which queries the index). The description implies this is for maintenance/rebuilding operations but doesn't specify triggers or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchB
Search for relevant document chunks using semantic similarity
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query to find relevant document chunks | |
| k | No | Number of top results to return (default: 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'semantic similarity' as the search method but doesn't describe other behavioral traits such as performance characteristics (e.g., speed, accuracy), error handling, or what happens if no results are found. For a search tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Search for relevant document chunks using semantic similarity.' It is front-loaded with the core purpose, has zero wasted words, and is appropriately sized for the tool's complexity. Every part of the sentence earns its place by specifying key details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose and method but lacks details on usage guidelines, behavioral traits, and output expectations. Without an output schema, the description doesn't explain return values, leaving the agent to infer results from the context. It meets the minimum viable standard but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with clear documentation for both parameters ('query' and 'k'). The description adds no additional meaning beyond the schema, such as explaining how 'semantic similarity' applies to the query or detailing result formats. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search for relevant document chunks using semantic similarity.' It specifies the action (search), the target (document chunks), and the method (semantic similarity). However, it doesn't explicitly differentiate from sibling tools like 'get_chunk' (which might retrieve a specific chunk) or 'refresh_index' (which might update search indices).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'get_chunk' (for direct retrieval), 'ingest_docs' (for adding documents), or 'refresh_index' (for updating indices), nor does it specify contexts or prerequisites for use. The agent must infer usage from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v1.0.0- First observed
get_chunk - First observed
ingest_docs - First observed
refresh_index - First observed
search
TDQS
Each tool has a clearly distinct purpose with no overlap: get_chunk retrieves a specific chunk, ingest_docs re-ingests documents, refresh_index clears and rebuilds the index, and search performs semantic similarity queries. The descriptions make it easy to differentiate between retrieval, ingestion, index management, and search operations.
All tool names follow a consistent verb_noun pattern (e.g., get_chunk, ingest_docs, refresh_index, search), using snake_case throughout. The naming is predictable and readable, with no deviations in style or convention across the set.
With 4 tools, the count is reasonable for a RAG server's core operations, covering ingestion, indexing, retrieval, and search. It feels slightly thin but well-scoped, as each tool earns its place without bloat, though additional utilities like document deletion or status checks might be considered minor gaps.
The toolset covers essential RAG workflows: ingestion (ingest_docs), index management (refresh_index), retrieval (get_chunk), and search (search). Minor gaps exist, such as no explicit update or delete operations for documents or chunks, but agents can work around this by re-ingesting or refreshing the index as needed.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Ingest, manage, and retrieve documents for RAG-powered AI applications
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
1Search your AI chat history (ChatGPT, Claude, Codex) from any MCP client. Remote, private, read-only
Self-hosted AI-native knowledge workspace with hybrid search, GraphRAG, and MCP.
Related MCP Servers
- AlicenseBqualityDmaintenanceA complete MCP server for Retrieval-Augmented Generation with file management and vector memory for agents. Supports multiple document formats (PDF, DOCX, TXT, MD, CSV, JSON) with semantic search using Hugging Face embeddings and ChromaDB for efficient vector storage.11121MIT
- AlicenseNot gradedqualityDmaintenanceProvides retrieval-augmented generation (RAG) capabilities by ingesting various document formats into a persistent ChromaDB vector store. It enables semantic search and retrieval using either OpenAI or Ollama embeddings for processing local files, directories, and URLs.1MIT
- AlicenseAqualityDmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.316MIT
- FlicenseNot gradedqualityCmaintenanceMCP server that provides 8 local RAG tools using LlamaIndex and Ollama, enabling AI-powered document querying, summarization, analysis, and comparison over PDFs, DOCX, XLSX, and CSV files.-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/LuizDoPc/mcp-rag'
If you have feedback or need assistance with the MCP directory API, please join our Discord server