Skip to main content
Glama
Ray0907
by Ray0907

wiki-compiler

A CLI + MCP server that compiles web sources into a structured, growing knowledge base — Obsidian-compatible, local-first, schema-driven.

"There is room here for an incredible new product instead of a hacky collection of scripts." — Andrej Karpathy

Most tools give your AI access to your notes. This one gives your AI a compiled knowledge structure — backlinked wiki pages built from raw sources, organized by a schema you define.


How it works

wiki add <url>
    │
    ▼
fetch page → strip HTML → send to Claude
    │
    ▼
Claude extracts: title, summary, key_concepts, categories, tags
    │
    ▼
vault/agent/{slug}.md  ← Obsidian-compatible markdown
    │
    ▼
wiki query "what do I know about X?"
    │
    ▼
Claude reads all vault pages → synthesized answer

Dual vault architecture (kepano's contamination insight):

  • vault/agent/ — compiler output, messy, experimental

  • vault/personal/ — your curated notes, never touched by the tool


Related MCP server: LLM Wiki MCP

Install

git clone https://github.com/Ray0907/wiki-compiler
cd wiki-compiler
uv sync
export ANTHROPIC_API_KEY=sk-ant-...

Usage

# Compile a source into your vault
uv run wiki.py add https://github.com/kepano/obsidian-skills

# Written: vault/agent/kepanoobsidian-skills-agent-skills-for-obsidian.md

# Ask a question against everything you've compiled
uv run wiki.py query "What is the Obsidian CLI used for?"

Then open vault/ as an Obsidian vault — every [[wikilink]] in the compiled pages becomes a node in your knowledge graph.


Output format

Each compiled page:

---
title: kepano/obsidian-skills
url: https://github.com/kepano/obsidian-skills
summary: A collection of agent skill files that teach AI agents how to read
  and write Obsidian vaults using Markdown, Bases, JSON Canvas, and CLI.
categories:
  - AI
  - Tools
tags:
  - obsidian
  - agents
  - mcp
date_compiled: '2026-04-03'
---
## Summary
A collection of agent skill files...

## Key Concepts
- [[Obsidian CLI]]
- [[Agent Skills]]
- [[JSON Canvas]]
- [[Model Context Protocol]]
- [[Vault Structure]]

Schema

schema.yaml controls what the compiler produces. Edit it to match your domain:

categories:
  - AI
  - Tools
  - Research
  - Knowledge Management
  - Agent Architecture

backlinks:
  min_concepts: 2
  max_concepts: 8
  format: "[[{concept}]]"

output:
  filename_pattern: "{slug}.md"
  frontmatter_fields:
    - title
    - url
    - summary
    - categories
    - tags
    - date_compiled

Schema is re-read on every run — no restart needed.


MCP Server

Connect to Claude Desktop, Cursor, or any MCP client:

uv run mcp_server.py

Two tools exposed:

  • add_document(url) — compile a URL into the vault

  • search_wiki(query) — query the vault

Claude Desktop config (claude_desktop_config.json):

{
  "mcpServers": {
    "wiki": {
      "command": "/path/to/wiki-compiler/.venv/bin/python",
      "args": ["/path/to/wiki-compiler/mcp_server.py"],
      "cwd": "/path/to/wiki-compiler",
      "env": {
        "ANTHROPIC_API_KEY": "sk-ant-..."
      }
    }
  }
}

Tests

uv run pytest tests/ -v
# 16 passed in 0.45s

Why not an Obsidian plugin?

Obsidian plugins can't run heavy LLM pipelines, are tied to one app, and the Obsidian team is already building native AI. The value here is the compiler — the pipeline that turns raw sources into a coherent, schema-consistent knowledge structure. Obsidian is just a convenient viewer for the output.

The MCP server makes the compiled vault available to any LLM client. As models improve, the compilation gets better — the schema and accumulated knowledge stay yours.


Roadmap

  • wiki compile — batch mode, multiple sources at once

  • Human review gate — promote pages from vault/agent/ to vault/personal/

  • wiki report — autonomous mode: one query spawns a team of agents, builds an ephemeral wiki, returns a full report

  • Obsidian CLI integration — use backlink graph for smarter query routing

Available Tools

3 tools
add_documentB

Fetch and compile a URL into the wiki vault.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'fetch and compile' but does not disclose side effects such as creating/modifying files in the vault, network requests, auth requirements, or failure modes. It is vague about what 'compile' entails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler. Every word contributes to conveying the tool's core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple tool with one parameter, but it still lacks important context: no mention of side effects, when to use it, or what happens on success/failure. The description is too minimal for an agent to safely invoke it without further assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage, but the description does indicate that the 'url' parameter is the target URL to fetch and compile. However, it adds no detail about URL format, validity, or edge cases. It provides essential context but not much beyond the parameter property name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Fetch and compile a URL into the wiki vault' uses a specific verb ('fetch and compile') and names the resource (URL into the wiki vault), clearly distinguishing it from sibling tools like search_wiki and lint_wiki. It clearly states what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description does not mention prerequisites, exclusions, or situations where search_wiki or lint_wiki would be more appropriate. Usage context is only implied by the verb and name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lint_wikiA

Audit the wiki vault for contradictions, orphans, and missing cross-references.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full burden. It accurately describes the audit's focus, but does not disclose whether the tool is read-only, returns only a report, or if any side effects occur. 'Audit' hints at non-modifying behavior, but explicit statements about output or side effects would improve transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that front-loads the verb and resource. Every word adds value, specifying the exact types of audit findings without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless lint/audit tool, the description is quite complete. It communicates the tool's purpose and scope. While it does not describe the output format, the absence of an output schema and the simplicity of the tool make this acceptable; the description still gives a clear picture of behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so schema coverage is effectively 100%. The description does not need to explain parameter behavior. The tool's purpose as an audit with no inputs is clear, matching the baseline for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Audit') and resource ('the wiki vault'), and narrows the scope with the types of findings ('contradictions, orphans, and missing cross-references'). This distinctly differentiates it from siblings add_document and search_wiki, which are separate operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: when you need to audit the wiki for quality issues. While it does not explicitly name alternatives or exclusions, the context is clear and the sibling tools (add_document, search_wiki) are obviously different, providing indirect when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_wikiC

Answer a question using the local wiki vault.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. It does not mention return format, whether it searches full text, whether it synthesizes answers, limitations, or any side effects. This is a significant gap for a tool that 'answers questions'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence), which is concise, but the brevity sacrifices clarity. It does not provide enough information to be adequately helpful, so it scores mid-range.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a single parameter, no annotations, and no output schema, so the description must compensate. It only says it answers questions from a local wiki, giving no context about result format, performance, or typical use cases. This is minimally viable but lacking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero description coverage and only one parameter 'query'. The description does not add any semantic detail about the query beyond the vague implication that it is a question, failing to explain expected format, language, or scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('answer a question') tied to a resource ('the local wiki vault'), and the name 'search_wiki' supports this. It distinguishes from sibling tools 'add_document' and 'lint_wiki' by implying a retrieval/QA function, though 'answer' is somewhat ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives, no mention of prerequisites or exclusions, and does not reference sibling tools. It merely states what the tool does.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.1.0
    • First observedadd_document
    • First observedlint_wiki
    • First observedsearch_wiki

TDQS

B3.3/5.0
Disambiguation5/5

Each tool targets a distinct operation: adding documents, searching the wiki, and auditing it. There is no overlap in purpose, and the descriptions clearly differentiate ingestion, query, and maintenance tasks.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern: add_document, search_wiki, lint_wiki. This makes the set predictable and easy to navigate.

Tool Count4/5

With only 3 tools, the server is at the low end of the typical range. The tools are focused and each serves a clear purpose, but the set is slightly sparse for a wiki compiler.

Completeness3/5

The core operations of adding, searching, and linting are covered, but there are no update, delete, or list document tools. This creates notable gaps in the document lifecycle and forces agents to work around them.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for persistent, compounding markdown wikis maintained by LLMs. Enables incremental knowledge base building with interlinked pages, search, and raw source management.
    33
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Compiles an Obsidian knowledge graph into a queryable MCP server, allowing AI agents to search concepts, get nodes, and retrieve related information.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to fetch webpages, convert them to Markdown, index into SQLite FTS5, and query the knowledge base through MCP tools.
    -

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Ray0907/wiki-compiler'

If you have feedback or need assistance with the MCP directory API, please join our Discord server