Knowledge Retrieval Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Knowledge Retrieval ServerWhat's our company's return policy for the AI platform?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Knowledge Retrieval Server
BM25 기반 문서 검색 및 검색을 위한 MCP(Model Context Protocol) 서버입니다.
참고한 토스결제연동 MCP 기술블로그. https://toss.tech/article/tosspayments-mcp
🚀 즉시 시작하기
🎯 초간단 설치 (권장)
# macOS/Linux
./run.sh
# Windows
run.bat이 스크립트들이 자동으로 처리합니다:
✅ Node.js 버전 확인
✅ 의존성 설치 (
npm install)✅ 프로젝트 빌드 (
npm run build)✅ 설정 파일 생성 (
config.json)✅ 예시 문서 생성 (
docs/폴더)✅ Claude Desktop 설정 가이드 출력
✅ MCP 서버 실행
수동 설치
# 단계별 설치
npm install && npm run build && cp config.example.json config.json
# 서버 실행
npm start개발 모드
npm run devRelated MCP server: DocuMCP
📋 기본 설정
config.json (자동 생성됨)
{
"serverName": "knowledge-retrieval",
"serverVersion": "1.0.0",
"documentSource": {
"type": "local",
"basePath": "./docs",
"domains": [
{
"name": "company",
"path": "company",
"category": "회사정보"
},
{
"name": "customer",
"path": "customer",
"category": "고객서비스"
},
{
"name": "product",
"path": "product",
"category": "제품정보"
},
{
"name": "technical",
"path": "technical",
"category": "기술문서"
}
]
},
"bm25": {
"k1": 1.2,
"b": 0.75
},
"chunk": {
"minWords": 30,
"contextWindowSize": 1
},
"logLevel": "info"
}주요 설정 항목
documentSource.basePath: 문서 파일들이 위치한 기본 경로
domains: 검색할 도메인들의 설정
bm25.k1: BM25 알고리즘의 term frequency saturation 파라미터 (기본값: 1.2)
bm25.b: BM25 알고리즘의 field length normalization 파라미터 (기본값: 0.75)
chunk.minWords: 청크의 최소 단어 수 (기본값: 30)
🔧 Claude Desktop 연동
설정 파일 위치
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%/Claude/claude_desktop_config.json
설정 내용 (절대 경로)
{
"mcpServers": {
"knowledge-retrieval": {
"command": "node",
"args": ["<프로젝트_경로>/dist/index.js"],
"env": {
"NODE_ENV": "production"
}
}
}
}중요: <프로젝트_경로>를 실제 프로젝트 폴더의 절대 경로로 바꾸세요!
권장 설정 (작업 디렉토리 지정)
{
"mcpServers": {
"knowledge-retrieval": {
"command": "npm",
"args": ["start"],
"cwd": "<프로젝트_경로>"
}
}
}📁 문서 구조
문서는 다음과 같은 구조로 구성되어야 합니다:
docs/
├── company/ # 회사 정보
│ ├── about.md
│ └── team.md
├── customer/ # 고객 서비스
│ ├── support.md
│ └── sla.md
├── product/ # 제품 정보
│ ├── ai-platform.md
│ └── web-app.md
└── technical/ # 기술 문서
├── api-guide.md
└── deployment.md지원 파일 형식
.md(Markdown).mdx(MDX).markdown
🛠 사용 가능한 MCP 도구들
1. search-documents
문서 검색을 수행합니다.
파라미터:
keywords: 검색할 키워드 배열maxResults: 최대 결과 수 (기본값: 10)domain: 특정 도메인으로 검색 제한 (선택사항)
예시:
// Claude Desktop에서 사용할 때
"AI 플랫폼의 가격 정책을 알려줘"2. get-document-by-id
특정 문서 ID로 전체 문서를 가져옵니다.
파라미터:
documentId: 문서 ID
3. list-domains
사용 가능한 모든 도메인과 문서 수를 조회합니다.
4. get-chunk-with-context
특정 청크와 그 주변 컨텍스트를 가져옵니다.
파라미터:
chunkId: 청크 IDcontextSize: 컨텍스트 윈도우 크기 (선택사항)
🧪 테스트 및 검증
1. 서버 작동 확인
npm run dev✅ 성공시 출력 예시:
Initializing knowledge-retrieval v1.0.0...
Loaded 8 documents
Initialized repository with 36 chunks from 8 documents
MCP server started successfully2. Claude Desktop에서 즉시 테스트
Claude Desktop 재시작 후 다음 질문들로 테스트:
우리 회사의 비전과 미션이 뭐야?
AI 플랫폼의 가격 정책을 알려줘
API 인증 방법을 설명해줘3. 빠른 문제 해결
문제 | 해결 방법 |
서버 시작 실패 |
|
문서 로드 실패 |
|
Claude Desktop 연결 실패 | 설정 파일 경로 확인 후 Claude Desktop 재시작 |
📊 성능 최적화
BM25 파라미터 튜닝
k1 값 증가: 단어 빈도의 영향 증가 (1.2 → 2.0)
b 값 조정: 문서 길이 정규화 강도 (0.75 → 0.5)
청크 크기 최적화
minWords 증가: 더 큰 컨텍스트, 느린 검색
minWords 감소: 정확한 매칭, 빠른 검색
🔒 보안 고려사항
파일 권한: 문서 디렉토리에 적절한 읽기 권한 설정
환경 변수: 민감한 설정은 환경 변수로 관리
네트워크: 필요시 방화벽 규칙 설정
📝 환경 변수 설정
export MCP_SERVER_NAME="my-knowledge-server"
export DOCS_BASE_PATH="./my-docs"
export BM25_K1="1.5"
export BM25_B="0.8"
export CHUNK_MIN_WORDS="50"
export LOG_LEVEL="debug"🆘 문제 해결
문제 발생 시 확인 순서:
로그 확인:
npm run dev출력 메시지설정 파일:
config.json문법 오류 확인문서 폴더:
docs/디렉토리와.md파일 확인Claude Desktop: 설정 파일 경로 및 재시작
💡 핵심 요약
즉시 사용을 위한 체크리스트
자동 설치 사용시:
./run.sh(또는run.bat) 실행스크립트가 출력하는 Claude Desktop 설정 복사
Claude Desktop 재시작
테스트 질문으로 작동 확인
수동 설치 사용시:
npm install && npm run build && cp config.example.json config.jsondocs/폴더에 마크다운 파일 추가Claude Desktop 설정 파일에 프로젝트 경로 지정
Claude Desktop 재시작
테스트 질문으로 작동 확인
주요 명령어
개발:
npm run dev빌드:
npm run build실행:
npm start테스트:
npm test
MIT 라이선스 | 개발 중에는 npm run dev 사용 권장
Available Tools
4 toolsget-chunk-with-contextC
Get specific chunk with surrounding context.
| Name | Required | Description | Default |
|---|---|---|---|
| documentId | Yes | Document ID | |
| chunkId | Yes | Chunk ID | |
| windowSize | No | Context window size (default: 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves a chunk with context, but doesn't describe what 'surrounding context' entails, whether it's read-only, potential errors, or response format. This is a significant gap for a tool with no annotation coverage, as it leaves key behavioral traits unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded with a single sentence that directly states the tool's function. There is no wasted language or unnecessary elaboration, making it efficient for quick comprehension. Every word earns its place by conveying the core purpose without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of retrieving chunks with context, no annotations, and no output schema, the description is incomplete. It doesn't explain what a chunk is, how context is provided, or what the return value includes. For a tool with 3 parameters and no structured support, this leaves too many gaps for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal meaning beyond the input schema, which has 100% coverage. It implies parameters for documentId, chunkId, and windowSize but doesn't explain their relationships or semantics (e.g., how chunkId relates to documentId, what windowSize units are). With high schema coverage, the baseline is 3, and the description doesn't significantly enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool's purpose ('Get specific chunk with surrounding context'), which is clear but vague. It specifies the verb 'Get' and resource 'chunk with surrounding context', but doesn't distinguish from siblings like 'get-document-by-id' or explain what a 'chunk' is in this context. The purpose is understandable but lacks specificity about what constitutes a chunk versus a document.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention siblings like 'get-document-by-id' or 'search-documents', nor does it specify prerequisites or exclusions. Usage is implied from the name and description but not explicitly stated, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-document-by-idC
Retrieve full document by ID.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Document ID to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states this is a retrieval operation but lacks details on permissions, rate limits, error handling, or response format. For a read tool with zero annotation coverage, this leaves significant gaps in understanding how it behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and wastes no space, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what a 'full document' entails (e.g., content, metadata), potential errors, or usage context. For a retrieval tool with no structured behavioral data, more detail is needed to fully inform an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'id' parameter clearly documented as 'Document ID to retrieve'. The description adds no additional meaning beyond this, such as format constraints or examples. With high schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Retrieve') and resource ('full document by ID'), making the purpose immediately understandable. However, it doesn't differentiate this from sibling tools like 'get-chunk-with-context' or 'search-documents', which likely also retrieve documents but with different approaches or scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a specific document ID), exclusions (e.g., not for partial documents), or comparisons to siblings like 'search-documents' for broader queries or 'get-chunk-with-context' for contextual retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list-domainsB
List all available domains and their document counts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool lists domains and their document counts, which implies a read-only operation, but doesn't address potential limitations like pagination, rate limits, authentication requirements, or whether the list is comprehensive versus filtered. This leaves significant gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without any wasted words. It's front-loaded with the core action ('list all available domains') and adds only essential additional context ('and their document counts'). Every part of the sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate but has clear gaps. It explains what the tool does but lacks context about when to use it, behavioral constraints, or output format details. For a list operation with no structured guidance, this is the minimum viable description—it covers the basics but leaves important aspects unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so the schema fully documents that no inputs are required. The description doesn't need to add parameter information, and it appropriately doesn't mention any parameters. A baseline of 4 is applied for tools with zero parameters, as the description correctly focuses on the tool's purpose rather than unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the resource 'domains', specifying what the tool does. It adds useful context about including 'document counts' in the output. However, it doesn't explicitly differentiate from sibling tools like 'search-documents' or 'get-document-by-id', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'search-documents' or 'get-document-by-id'. It doesn't mention prerequisites, limitations, or specific contexts where listing all domains is appropriate versus searching for specific documents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search-documentsC
Search documents using BM25 algorithm. Takes keyword arrays and returns relevant document chunks.
| Name | Required | Description | Default |
|---|---|---|---|
| keywords | Yes | Array of keywords to search for (e.g., ["payment", "API", "authentication"]) | |
| domain | No | Domain to search in (optional, e.g., "company", "customer") | |
| topN | No | Maximum number of results to return (default: 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the BM25 algorithm and that it returns document chunks, but lacks critical details: it doesn't specify if this is a read-only operation (implied but not stated), whether it has rate limits, authentication needs, or how results are ranked/ordered. For a search tool with no annotations, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and front-loaded: two sentences that directly state the action ('Search documents using BM25 algorithm'), inputs ('Takes keyword arrays'), and outputs ('returns relevant document chunks'). Every word earns its place with no redundancy or fluff, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a search tool with 3 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain the return format (e.g., structure of document chunks), error handling, or performance characteristics. While the purpose is clear, the lack of behavioral details and output information makes it inadequate for full contextual understanding, especially without annotations to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters ('keywords', 'domain', 'topN') with descriptions and defaults. The description adds minimal value beyond the schema—it mentions 'keyword arrays' and 'returns relevant document chunks', but doesn't explain parameter interactions or provide additional context like format examples beyond what's in the schema. This meets the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search documents using BM25 algorithm' specifies the verb (search) and resource (documents), and 'returns relevant document chunks' clarifies the output. It distinguishes from siblings like 'get-document-by-id' (retrieval by ID) and 'list-domains' (listing domains), though not explicitly. However, it doesn't fully differentiate from 'get-chunk-with-context' (which might retrieve specific chunks), keeping it at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose this over 'get-chunk-with-context' for contextual retrieval or 'list-domains' for domain exploration. There's no context about prerequisites, exclusions, or typical use cases, leaving the agent to infer usage from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
- First observed
get-chunk-with-context - First observed
get-document-by-id - First observed
list-domains - First observed
search-documents
TDQS
Each tool has a clearly distinct purpose with no overlap: get-chunk-with-context retrieves specific text segments, get-document-by-id fetches entire documents, list-domains provides metadata, and search-documents performs keyword-based searches. The descriptions reinforce these unique functions, eliminating any ambiguity.
The tools follow a consistent verb_noun pattern (e.g., get-chunk-with-context, get-document-by-id, list-domains, search-documents), with all using hyphens for readability. The minor deviation is that 'get-chunk-with-context' includes a prepositional phrase, but overall naming remains predictable and clear.
With 4 tools, the server is well-scoped for knowledge retrieval, covering core operations like listing domains, retrieving documents/chunks, and searching. This count is efficient and avoids bloat, with each tool serving a distinct and necessary function in the domain.
The tool set covers essential retrieval workflows: metadata listing, full document access, chunk retrieval, and search. A minor gap is the lack of update or delete operations, but this aligns with a retrieval-focused server, and agents can work effectively with the provided read-only tools.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Augments MCP Server - A comprehensive framework documentation provider for Claude Code
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
MCP server giving Claude AI access to 22+ NYC public-record databases for real estate due diligence
The AWS Knowledge MCP server is a fully managed remote Model Context Protocol server that provides real-time access to official AWS content in an LLM-compatible format. It offers structured access to AWS documentation, code samples, blog posts, What's New announcements, Well-Architected best practices, and regional availability information for AWS APIs and CloudFormation resources. Key capabilities include searching and reading documentation in markdown format, getting content recommendations, listing AWS regions, and checking regional availability for services and features.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that integrates with Claude to provide smart documentation search capabilities across multiple AI/ML libraries, allowing users to retrieve and process technical information through natural language queries.-
- FlicenseNot gradedqualityNot gradedmaintenanceAn MCP server that enables Claude to generate, search, and manage documentation for codebases using vector embeddings and semantic search, providing tools for creating user guides, technical documentation, code explanations, and architectural diagrams.6-
- AlicenseBqualityBmaintenanceAn MCP server that provides searchable local storage for Claude conversation history, featuring automatic topic extraction and weekly insight summaries. It enables Claude to retrieve context from past sessions through full-text search and organized file storage.103MIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that enables Claude Desktop to search and read local documents via full-text and fuzzy search, providing direct access to indexed files without chunking.MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/cskwork/keyword-rag-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server