Skip to main content
Glama
cskwork

Knowledge Retrieval Server

by cskwork

MCP Knowledge Retrieval Server

BM25 기반 문서 검색 및 검색을 위한 MCP(Model Context Protocol) 서버입니다.

🚀 즉시 시작하기

🎯 초간단 설치 (권장)

# macOS/Linux
./run.sh

# Windows
run.bat

이 스크립트들이 자동으로 처리합니다:

  • ✅ Node.js 버전 확인

  • ✅ 의존성 설치 (npm install)

  • ✅ 프로젝트 빌드 (npm run build)

  • ✅ 설정 파일 생성 (config.json)

  • ✅ 예시 문서 생성 (docs/ 폴더)

  • ✅ Claude Desktop 설정 가이드 출력

  • ✅ MCP 서버 실행

수동 설치

# 단계별 설치
npm install && npm run build && cp config.example.json config.json

# 서버 실행
npm start

개발 모드

npm run dev

Related MCP server: DocuMCP

📋 기본 설정

config.json (자동 생성됨)

{
  "serverName": "knowledge-retrieval",  
  "serverVersion": "1.0.0",
  "documentSource": {
    "type": "local",
    "basePath": "./docs",
    "domains": [
      {
        "name": "company",
        "path": "company",
        "category": "회사정보"
      },
      {
        "name": "customer", 
        "path": "customer",
        "category": "고객서비스"
      },
      {
        "name": "product",
        "path": "product", 
        "category": "제품정보"
      },
      {
        "name": "technical",
        "path": "technical",
        "category": "기술문서"
      }
    ]
  },
  "bm25": {
    "k1": 1.2,
    "b": 0.75
  },
  "chunk": {
    "minWords": 30,
    "contextWindowSize": 1
  },
  "logLevel": "info"
}

주요 설정 항목

  • documentSource.basePath: 문서 파일들이 위치한 기본 경로

  • domains: 검색할 도메인들의 설정

  • bm25.k1: BM25 알고리즘의 term frequency saturation 파라미터 (기본값: 1.2)

  • bm25.b: BM25 알고리즘의 field length normalization 파라미터 (기본값: 0.75)

  • chunk.minWords: 청크의 최소 단어 수 (기본값: 30)

🔧 Claude Desktop 연동

설정 파일 위치

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%/Claude/claude_desktop_config.json

설정 내용 (절대 경로)

{
  "mcpServers": {
    "knowledge-retrieval": {
      "command": "node",
      "args": ["<프로젝트_경로>/dist/index.js"],
      "env": {
        "NODE_ENV": "production"
      }
    }
  }
}

중요: <프로젝트_경로>를 실제 프로젝트 폴더의 절대 경로로 바꾸세요!

권장 설정 (작업 디렉토리 지정)

{
  "mcpServers": {
    "knowledge-retrieval": {
      "command": "npm",
      "args": ["start"],
      "cwd": "<프로젝트_경로>"
    }
  }
}

📁 문서 구조

문서는 다음과 같은 구조로 구성되어야 합니다:

docs/
├── company/           # 회사 정보
│   ├── about.md
│   └── team.md
├── customer/          # 고객 서비스
│   ├── support.md
│   └── sla.md
├── product/           # 제품 정보
│   ├── ai-platform.md
│   └── web-app.md
└── technical/         # 기술 문서
    ├── api-guide.md
    └── deployment.md

지원 파일 형식

  • .md (Markdown)

  • .mdx (MDX)

  • .markdown

🛠 사용 가능한 MCP 도구들

1. search-documents

문서 검색을 수행합니다.

파라미터:

  • keywords: 검색할 키워드 배열

  • maxResults: 최대 결과 수 (기본값: 10)

  • domain: 특정 도메인으로 검색 제한 (선택사항)

예시:

// Claude Desktop에서 사용할 때
"AI 플랫폼의 가격 정책을 알려줘"

2. get-document-by-id

특정 문서 ID로 전체 문서를 가져옵니다.

파라미터:

  • documentId: 문서 ID

3. list-domains

사용 가능한 모든 도메인과 문서 수를 조회합니다.

4. get-chunk-with-context

특정 청크와 그 주변 컨텍스트를 가져옵니다.

파라미터:

  • chunkId: 청크 ID

  • contextSize: 컨텍스트 윈도우 크기 (선택사항)

🧪 테스트 및 검증

1. 서버 작동 확인

npm run dev

✅ 성공시 출력 예시:

Initializing knowledge-retrieval v1.0.0...
Loaded 8 documents
Initialized repository with 36 chunks from 8 documents
MCP server started successfully

2. Claude Desktop에서 즉시 테스트

Claude Desktop 재시작 후 다음 질문들로 테스트:

우리 회사의 비전과 미션이 뭐야?
AI 플랫폼의 가격 정책을 알려줘
API 인증 방법을 설명해줘

3. 빠른 문제 해결

문제

해결 방법

서버 시작 실패

npm install && npm run build

문서 로드 실패

docs/ 폴더와 .md 파일 확인

Claude Desktop 연결 실패

설정 파일 경로 확인 후 Claude Desktop 재시작

📊 성능 최적화

BM25 파라미터 튜닝

  • k1 값 증가: 단어 빈도의 영향 증가 (1.2 → 2.0)

  • b 값 조정: 문서 길이 정규화 강도 (0.75 → 0.5)

청크 크기 최적화

  • minWords 증가: 더 큰 컨텍스트, 느린 검색

  • minWords 감소: 정확한 매칭, 빠른 검색

🔒 보안 고려사항

  1. 파일 권한: 문서 디렉토리에 적절한 읽기 권한 설정

  2. 환경 변수: 민감한 설정은 환경 변수로 관리

  3. 네트워크: 필요시 방화벽 규칙 설정

📝 환경 변수 설정

export MCP_SERVER_NAME="my-knowledge-server"
export DOCS_BASE_PATH="./my-docs"
export BM25_K1="1.5"
export BM25_B="0.8"
export CHUNK_MIN_WORDS="50"
export LOG_LEVEL="debug"

🆘 문제 해결

문제 발생 시 확인 순서:

  1. 로그 확인: npm run dev 출력 메시지

  2. 설정 파일: config.json 문법 오류 확인

  3. 문서 폴더: docs/ 디렉토리와 .md 파일 확인

  4. Claude Desktop: 설정 파일 경로 및 재시작

💡 핵심 요약

즉시 사용을 위한 체크리스트

자동 설치 사용시:

  • ./run.sh (또는 run.bat) 실행

  • 스크립트가 출력하는 Claude Desktop 설정 복사

  • Claude Desktop 재시작

  • 테스트 질문으로 작동 확인

수동 설치 사용시:

  • npm install && npm run build && cp config.example.json config.json

  • docs/ 폴더에 마크다운 파일 추가

  • Claude Desktop 설정 파일에 프로젝트 경로 지정

  • Claude Desktop 재시작

  • 테스트 질문으로 작동 확인

주요 명령어

  • 개발: npm run dev

  • 빌드: npm run build

  • 실행: npm start

  • 테스트: npm test


MIT 라이선스 | 개발 중에는 npm run dev 사용 권장

Available Tools

4 tools
get-chunk-with-contextC

Get specific chunk with surrounding context.

ParametersJSON Schema
NameRequiredDescriptionDefault
documentIdYesDocument ID
chunkIdYesChunk ID
windowSizeNoContext window size (default: 1)

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves a chunk with context, but doesn't describe what 'surrounding context' entails, whether it's read-only, potential errors, or response format. This is a significant gap for a tool with no annotation coverage, as it leaves key behavioral traits unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded with a single sentence that directly states the tool's function. There is no wasted language or unnecessary elaboration, making it efficient for quick comprehension. Every word earns its place by conveying the core purpose without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of retrieving chunks with context, no annotations, and no output schema, the description is incomplete. It doesn't explain what a chunk is, how context is provided, or what the return value includes. For a tool with 3 parameters and no structured support, this leaves too many gaps for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal meaning beyond the input schema, which has 100% coverage. It implies parameters for documentId, chunkId, and windowSize but doesn't explain their relationships or semantics (e.g., how chunkId relates to documentId, what windowSize units are). With high schema coverage, the baseline is 3, and the description doesn't significantly enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool's purpose ('Get specific chunk with surrounding context'), which is clear but vague. It specifies the verb 'Get' and resource 'chunk with surrounding context', but doesn't distinguish from siblings like 'get-document-by-id' or explain what a 'chunk' is in this context. The purpose is understandable but lacks specificity about what constitutes a chunk versus a document.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention siblings like 'get-document-by-id' or 'search-documents', nor does it specify prerequisites or exclusions. Usage is implied from the name and description but not explicitly stated, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-document-by-idC

Retrieve full document by ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesDocument ID to retrieve

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states this is a retrieval operation but lacks details on permissions, rate limits, error handling, or response format. For a read tool with zero annotation coverage, this leaves significant gaps in understanding how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and wastes no space, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what a 'full document' entails (e.g., content, metadata), potential errors, or usage context. For a retrieval tool with no structured behavioral data, more detail is needed to fully inform an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the 'id' parameter clearly documented as 'Document ID to retrieve'. The description adds no additional meaning beyond this, such as format constraints or examples. With high schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Retrieve') and resource ('full document by ID'), making the purpose immediately understandable. However, it doesn't differentiate this from sibling tools like 'get-chunk-with-context' or 'search-documents', which likely also retrieve documents but with different approaches or scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a specific document ID), exclusions (e.g., not for partial documents), or comparisons to siblings like 'search-documents' for broader queries or 'get-chunk-with-context' for contextual retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-domainsB

List all available domains and their document counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool lists domains and their document counts, which implies a read-only operation, but doesn't address potential limitations like pagination, rate limits, authentication requirements, or whether the list is comprehensive versus filtered. This leaves significant gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's function without any wasted words. It's front-loaded with the core action ('list all available domains') and adds only essential additional context ('and their document counts'). Every part of the sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, no output schema, no annotations), the description is adequate but has clear gaps. It explains what the tool does but lacks context about when to use it, behavioral constraints, or output format details. For a list operation with no structured guidance, this is the minimum viable description—it covers the basics but leaves important aspects unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so the schema fully documents that no inputs are required. The description doesn't need to add parameter information, and it appropriately doesn't mention any parameters. A baseline of 4 is applied for tools with zero parameters, as the description correctly focuses on the tool's purpose rather than unnecessary parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and the resource 'domains', specifying what the tool does. It adds useful context about including 'document counts' in the output. However, it doesn't explicitly differentiate from sibling tools like 'search-documents' or 'get-document-by-id', which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'search-documents' or 'get-document-by-id'. It doesn't mention prerequisites, limitations, or specific contexts where listing all domains is appropriate versus searching for specific documents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search-documentsC

Search documents using BM25 algorithm. Takes keyword arrays and returns relevant document chunks.

ParametersJSON Schema
NameRequiredDescriptionDefault
keywordsYesArray of keywords to search for (e.g., ["payment", "API", "authentication"])
domainNoDomain to search in (optional, e.g., "company", "customer")
topNNoMaximum number of results to return (default: 10)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the BM25 algorithm and that it returns document chunks, but lacks critical details: it doesn't specify if this is a read-only operation (implied but not stated), whether it has rate limits, authentication needs, or how results are ranked/ordered. For a search tool with no annotations, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and front-loaded: two sentences that directly state the action ('Search documents using BM25 algorithm'), inputs ('Takes keyword arrays'), and outputs ('returns relevant document chunks'). Every word earns its place with no redundancy or fluff, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with 3 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain the return format (e.g., structure of document chunks), error handling, or performance characteristics. While the purpose is clear, the lack of behavioral details and output information makes it inadequate for full contextual understanding, especially without annotations to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters ('keywords', 'domain', 'topN') with descriptions and defaults. The description adds minimal value beyond the schema—it mentions 'keyword arrays' and 'returns relevant document chunks', but doesn't explain parameter interactions or provide additional context like format examples beyond what's in the schema. This meets the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search documents using BM25 algorithm' specifies the verb (search) and resource (documents), and 'returns relevant document chunks' clarifies the output. It distinguishes from siblings like 'get-document-by-id' (retrieval by ID) and 'list-domains' (listing domains), though not explicitly. However, it doesn't fully differentiate from 'get-chunk-with-context' (which might retrieve specific chunks), keeping it at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose this over 'get-chunk-with-context' for contextual retrieval or 'list-domains' for domain exploration. There's no context about prerequisites, exclusions, or typical use cases, leaving the agent to infer usage from the purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updates
    • First observedget-chunk-with-context
    • First observedget-document-by-id
    • First observedlist-domains
    • First observedsearch-documents

TDQS

B3.3/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: get-chunk-with-context retrieves specific text segments, get-document-by-id fetches entire documents, list-domains provides metadata, and search-documents performs keyword-based searches. The descriptions reinforce these unique functions, eliminating any ambiguity.

Naming Consistency4/5

The tools follow a consistent verb_noun pattern (e.g., get-chunk-with-context, get-document-by-id, list-domains, search-documents), with all using hyphens for readability. The minor deviation is that 'get-chunk-with-context' includes a prepositional phrase, but overall naming remains predictable and clear.

Tool Count5/5

With 4 tools, the server is well-scoped for knowledge retrieval, covering core operations like listing domains, retrieving documents/chunks, and searching. This count is efficient and avoids bloat, with each tool serving a distinct and necessary function in the domain.

Completeness4/5

The tool set covers essential retrieval workflows: metadata listing, full document access, chunk retrieval, and search. A minor gap is the lack of update or delete operations, but this aligns with a retrieval-focused server, and agents can work effectively with the provided read-only tools.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that integrates with Claude to provide smart documentation search capabilities across multiple AI/ML libraries, allowing users to retrieve and process technical information through natural language queries.
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    An MCP server that enables Claude to generate, search, and manage documentation for codebases using vector embeddings and semantic search, providing tools for creating user guides, technical documentation, code explanations, and architectural diagrams.
    6
    -
  • A
    license
    B
    quality
    B
    maintenance
    An MCP server that provides searchable local storage for Claude conversation history, featuring automatic topic extraction and weekly insight summaries. It enables Claude to retrieve context from past sessions through full-text search and organized file storage.
    10
    3
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server that enables Claude Desktop to search and read local documents via full-text and fuzzy search, providing direct access to indexed files without chunking.
    MIT

Appeared in Searches

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/cskwork/keyword-rag-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server