Skip to main content
Glama

mcp-for-docs

GitHub Status Platform License Version

An MCP (Model Context Protocol) server that automatically downloads and converts documentation from various sources into organized markdown files.

Overview

mcp-for-docs is designed to crawl documentation websites, convert their content to markdown format, and organize them in a structured directory system. It can also generate condensed cheat sheets from the downloaded documentation.

Related MCP server: Documentation MCP Server

Features

  • πŸ•·οΈ Smart Documentation Crawler: Automatically crawls documentation sites with configurable depth

  • πŸ“ HTML to Markdown Conversion: Preserves code blocks, tables, and formatting

  • πŸ“ Automatic Categorization: Intelligently organizes docs into tools/APIs categories

  • πŸ“„ Cheat Sheet Generator: Creates condensed reference guides from documentation

  • πŸ” Smart Discovery System: Automatically detects existing documentation before crawling

  • πŸš€ Local-First: Uses existing downloaded docs when available

  • ⚑ Rate Limiting: Respects server limits and robots.txt

  • βœ… User Confirmation: Prevents accidental regeneration of existing content

  • βš™οΈ Comprehensive Configuration: JSON-based configuration with environment variable overrides

  • πŸ§ͺ Test Suite: 94 tests covering core functionality

Installation

Prerequisites

  • Node.js 18+

  • npm or yarn

  • Claude Desktop or Claude Code CLI

Setup

  1. Clone the repository:

git clone https://github.com/shayonpal/mcp-for-docs.git
cd mcp-for-docs
  1. Install dependencies:

npm install
  1. Build the project:

npm run build
  1. Add to your MCP configuration:

For Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "mcp-for-docs": {
      "command": "node",
      "args": ["/path/to/mcp-for-docs/dist/index.js"],
      "env": {}
    }
  }
}

For Claude Code CLI (~/.claude.json):

{
  "mcpServers": {
    "mcp-for-docs": {
      "command": "node",
      "args": ["/path/to/mcp-for-docs/dist/index.js"],
      "env": {}
    }
  }
}

Usage

Crawling Documentation

To download documentation from a website:

await crawl_documentation({
  url: "https://docs.n8n.io/",
  max_depth: 3,           // Optional, defaults to 3
  force_refresh: false    // Optional, set to true to regenerate existing docs
});

The tool will first check for existing documentation and show you what's already available. To regenerate existing content, use force_refresh: true.

The documentation will be saved to:

  • Tools: /Users/shayon/DevProjects/~meta/docs/tools/[tool-name]/

  • APIs: /Users/shayon/DevProjects/~meta/docs/apis/[api-name]/

Generating Cheat Sheets

To create a cheat sheet from documentation:

await generate_cheatsheet({
  url: "https://docs.anthropic.com/",
  use_local: true,          // Use local files if available (default)
  force_regenerate: false   // Optional, set to true to regenerate existing cheatsheets
});

Cheat sheets are saved to: /Users/shayon/DevProjects/~meta/docs/cheatsheets/

The tool will check for existing cheatsheets and show you what's already available. To regenerate existing content, use force_regenerate: true.

Listing Downloaded Documentation

To see what documentation is available locally:

await list_documentation({
  category: "all",  // Options: "tools", "apis", "all"
  include_stats: true
});

Supported Documentation Sites

The server has been tested with:

  • n8n documentation

  • Anthropic API docs

  • Obsidian Tasks plugin docs

  • Apple Swift documentation

Most documentation sites following standard patterns should work automatically.

Recent Updates

  • Configuration System (v0.4.0): Added comprehensive JSON-based configuration with environment variable support

  • Smart Discovery: Automatically finds and reports existing documentation before crawling

  • Improved Conversion: Fixed HTML to Markdown issues including table formatting and inline code preservation

  • Dynamic Categorization: Intelligent detection of tools vs APIs based on URL patterns and content analysis

  • Test Coverage: 94 tests passing with comprehensive unit and integration testing

For detailed changes, see CHANGELOG.md.

Configuration

Initial Setup

  1. Copy the example configuration:

cp config.example.json config.json
  1. Edit config.json and update the docsBasePath for your machine:

{
  "docsBasePath": "/Users/yourusername/path/to/docs"
}

Important: The config.json file is tracked in git. When you clone this repository on a different machine, you'll need to update the docsBasePath to match that machine's directory structure.

How Documentation Organization Works

The tool automatically organizes documentation based on content analysis:

  1. You provide a URL when calling the tool (e.g., https://docs.n8n.io)

  2. The categorizer analyzes the content and determines if it's:

    • tools/ - Software tools, applications, plugins

    • apis/ - API references, SDK documentation

  3. Documentation is saved to: {docsBasePath}/{category}/{tool-name}/

For example:

  • https://docs.n8n.io β†’ /Users/shayon/DevProjects/~meta/docs/tools/n8n/

  • https://docs.anthropic.com β†’ /Users/shayon/DevProjects/~meta/docs/apis/anthropic/

This happens automatically - you don't need to configure anything per-site!

Configuration Options

Setting

Description

Default

docsBasePath

Where to store all documentation

Required - no default

crawler.defaultMaxDepth

How many levels deep to crawl

3

crawler.defaultRateLimit

Requests per second

2

crawler.pageTimeout

Page load timeout (ms)

30000

crawler.userAgent

Browser identification

MCP-for-docs/1.0

cheatsheet.maxLength

Max characters in cheatsheet

10000

cheatsheet.filenameSuffix

Append to cheatsheet names

-Cheatsheet.md

Multi-Machine Setup

Since config.json is tracked in git:

  1. First machine: Set your docsBasePath and commit

  2. Other machines: After cloning, update docsBasePath to match that machine

  3. Use environment variable to override without changing the file:

    export DOCS_BASE_PATH="/different/path/on/this/machine"

Development

# Install dependencies
npm install

# Run in development mode
npm run dev

# Run tests
npm test

# Build for production
npm run build

# Lint code
npm run lint

Architecture

  • Crawler: Uses Playwright for JavaScript-rendered pages

  • Parser: Extracts content using configurable selectors

  • Converter: Turndown library with custom rules for markdown

  • Categorizer: Smart detection of tools vs APIs

  • Storage: Organized file system structure

Known Issues

  • URL Structure Preservation (#15): Currently flattens URL structure when saving docs

  • Large Documentation Sites (#14): No document limit for very large sites

  • GitHub Repository Docs (#9): Specialized crawler for GitHub repos not yet implemented

See all open issues for the complete roadmap.

Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Update CHANGELOG.md

  5. Submit a pull request

License

This project is licensed under the GPL 3.0 License - see the LICENSE file for details.

Acknowledgments

Available Tools

3 tools
crawl_documentationC

Crawl and download documentation from a website

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesDocumentation homepage URL
max_depthNoMaximum crawl depth
force_refreshNoForce refresh existing docs
rate_limitNoRequests per second
include_patternsNoURL patterns to include
exclude_patternsNoURL patterns to exclude

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'crawl and download' which implies network operations and data retrieval, but fails to disclose critical traits like authentication needs, rate limits (though hinted in schema), potential destructive effects (e.g., overwriting files), or output format. This leaves significant gaps for a tool with complex behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any fluff or redundant information. It's appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, network operations, no output schema, and no annotations), the description is insufficient. It doesn't explain what 'crawl' entails (e.g., recursive linking, HTML parsing), what 'download' produces (e.g., files, database entries), or behavioral constraints, leaving the agent with inadequate context for safe and effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, meaning all parameters are documented in the schema itself. The description adds no additional parameter semantics beyond what's in the schema (e.g., it doesn't explain how 'max_depth' relates to documentation structure or what 'include_patterns' typically look like). This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('crawl and download') and resource ('documentation from a website'), providing a specific verb+resource combination. However, it doesn't differentiate from sibling tools like 'list_documentation' or 'generate_cheatsheet', which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives like 'list_documentation' or 'generate_cheatsheet'. It lacks any context about prerequisites, when-not-to-use scenarios, or explicit alternatives, leaving the agent with minimal usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_cheatsheetC

Generate a cheat sheet from documentation

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesDocumentation URL
use_localNoUse local files if available
sectionsNoSpecific sections to include
output_formatNoOutput formatsingle
max_lengthNoMaximum characters
force_regenerateNoForce regenerate existing cheatsheets

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Generate a cheat sheet' which implies a creation/processing action, but it doesn't disclose key traits like whether this is a read-only or mutative operation, potential rate limits, authentication needs, or what happens with existing cheatsheets (e.g., caching behavior hinted by 'force_regenerate' in schema). The description is too vague to inform the agent adequately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence ('Generate a cheat sheet from documentation') that is front-loaded and wastes no words. It directly states the purpose without unnecessary elaboration, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (6 parameters, no annotations, no output schema), the description is incomplete. It doesn't address behavioral aspects like mutability or side effects, provide usage context relative to siblings, or explain output expectations (e.g., format or content of the cheat sheet). For a tool with multiple parameters and no structured output information, the description should do more to guide the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, meaning all parameters are documented in the schema itself (e.g., 'url' for documentation URL, 'sections' for specific sections). The description adds no additional meaning beyond the schema, such as explaining how parameters interact (e.g., 'use_local' with 'url') or typical use cases. With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but also doesn't detract.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Generate a cheat sheet') and the source ('from documentation'), which is specific and understandable. However, it doesn't differentiate from sibling tools like 'crawl_documentation' or 'list_documentation', which likely have different purposes (e.g., crawling vs. listing vs. generating summaries).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like the sibling tools. It lacks context such as prerequisites (e.g., needing accessible documentation), exclusions (e.g., not for raw data extraction), or comparisons to other tools, leaving the agent to infer usage based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_documentationC

List available documentation

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoCategory to listall
include_statsNoInclude file statistics

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states a read operation ('List'), implying it's non-destructive, but doesn't cover aspects like authentication needs, rate limits, pagination, or return format. For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. It's front-loaded and appropriately sized for a simple tool, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no output schema, no annotations), the description is incomplete. It lacks details on return values, behavioral traits, and differentiation from siblings. Without annotations or output schema, the agent has insufficient context to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear descriptions for both parameters (category and include_stats). The description doesn't add any meaning beyond what the schema provides, such as explaining the impact of include_stats or the categories. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'List available documentation' states the basic action (list) and resource (documentation), but it's vague about scope and format. It doesn't specify what 'available documentation' means or how it differs from sibling tools like crawl_documentation and generate_cheatsheet, which could involve similar documentation resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description doesn't mention sibling tools or contexts where list_documentation is preferred over crawl_documentation or generate_cheatsheet, leaving the agent to infer usage based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv1.0.0
    • First observedcrawl_documentation
    • First observedgenerate_cheatsheet
    • First observedlist_documentation

TDQS

B3.2/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: crawling/downloading documentation, generating a cheat sheet from it, and listing available documentation. There is no overlap in functionality, making it easy for an agent to select the right tool.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (crawl_documentation, generate_cheatsheet, list_documentation) with snake_case throughout. This predictability aids in tool discovery and usage.

Tool Count3/5

With only 3 tools, the set feels thin for a documentation management server. While the tools cover core operations, more comprehensive coverage (e.g., search, update, or delete documentation) would be expected for such a domain.

Completeness3/5

The tools cover listing, crawling, and generating cheat sheets, but there are notable gaps. Missing operations like searching within documentation, updating or deleting documentation, or managing documentation metadata limit the server's utility for full lifecycle management.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    A comprehensive, domain-agnostic documentation scraping and AI integration toolkit. Scrape any documentation website, create structured databases, and integrate with Claude Desktop via MCP (Model Context Protocol) for seamless AI-powered documentation assistance.
    5
    -
  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Crawls documentation websites and provides semantic search capabilities over the content through vector embeddings, enabling natural language queries of technical documentation.
    2
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables crawling and extracting clean content from documentation websites with optional LLM-powered analysis for intelligent summaries, code example extraction, and content classification.
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/shayonpal/mcp-for-docs'

If you have feedback or need assistance with the MCP directory API, please join our Discord server