docs-search-engine
Server Quality Checklist
Latest release: v0.1.0
- Disambiguation4/5
scrape_web and count_word_occurrences both involve scraping a URL, so there is slight overlap, but count_word_occurrences is clearly a specialized analysis and search_docs is distinct. An agent can generally select the right tool from the stated purpose.
Naming Consistency5/5All three tool names use a consistent verb_noun snake_case pattern (scrape_web, count_word_occurrences, search_docs), making the set predictable and easy to navigate.
Tool Count4/5Three tools is a small but reasonable surface for a focused toolkit; only one tool actually performs documentation search, which feels slightly lean for a server named docs-search-engine, but no tool is redundant.
Completeness4/5The core function of searching a GitHub repository's docs is implemented, but cache management and indexing of arbitrary scraped pages are not supported. These gaps are minor and can be worked around by passing a new zip_url or relying on the automatic index.
Average 4/5 across 3 of 3 tools scored.
See the Tool Scores section below for per-tool breakdowns.
- No community issues in the last 6 months
- 0 commits in the last 12 weeks
- No stable releases found
- No critical vulnerability alerts
- No high-severity vulnerability alerts
- No code scanning findings
- CI status not available
Add a LICENSE file by following GitHub's guide. Once GitHub recognizes the license, the system will automatically detect it within a few hours.
If the license does not appear after some time, you can manually trigger a new scan using the MCP server admin interface.
MCP servers without a LICENSE cannot be installed.
This repository includes a README.md file.
No tool usage detected in the last 30 days. Usage tracking helps demonstrate server value.
Tip: use the "Try in Browser" feature on the server page to seed initial usage.
Add a glama.json file to provide metadata about your server.
If you are the author, simply .
If the server belongs to an organization, first add
glama.jsonto the root of your repository:{ "$schema": "https://glama.ai/mcp/schemas/server.json", "maintainers": [ "your-github-username" ] }Then . Browse examples.
Add related servers to improve discoverability.
How to sync the server with GitHub?
Servers are automatically synced at least once per day, but you can also sync manually at any time to instantly update the server profile.
To manually sync the server, click the "Sync Server" button in the MCP server admin interface.
How is the quality score calculated?
The overall quality score combines two components: Tool Definition Quality (70%) and Server Coherence (30%).
Tool Definition Quality measures how well each tool describes itself to AI agents. Every tool is scored 1–5 across six dimensions: Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), and Contextual Completeness (10%). The server-level definition quality score is calculated as 60% mean TDQS + 40% minimum TDQS, so a single poorly described tool pulls the score down.
Server Coherence evaluates how well the tools work together as a set, scoring four dimensions equally: Disambiguation (can agents tell tools apart?), Naming Consistency, Tool Count Appropriateness, and Completeness (are there gaps in the tool surface?).
Tiers are derived from the overall score: A (≥3.5), B (≥3.0), C (≥2.0), D (≥1.0), F (<1.0). B and above is considered passing.
Tool Scores
- Behavior3/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It does disclose key behavior: it fetches external web content, returns markdown, uses Jina Reader, and requires http/https URLs. However, it does not mention errors, access limitations, size limits, or that this is a read-only network operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness3/5Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and includes clear Args/Returns sections, but it is somewhat redundant: 'Scrape the content of a web page using Jina Reader' and 'fetches the content of any web page... using the Jina Reader service' overlap. It could be tightened without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness4/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with an output schema, the description covers the essential contract: input URL must be http/https, output is the page content as markdown. It lacks failure-mode detail, but the tool is low-complexity and the return format is explicitly stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters4/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero description coverage, but the tool description compensates well for the single parameter: it identifies 'url' as the page to scrape and imposes the http/https constraint. This gives an agent meaningful information beyond the bare schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose4/5Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb and resource: scrape the content of a web page using Jina Reader, returning markdown. It does not explicitly compare against siblings such as search_docs, but the action and resource are clear enough to distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines3/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when raw web page content is needed, and states it can handle 'any web page'. It gives no explicit guidance about when not to use it or what alternatives (e.g., search_docs, count_word_occurrences) are better suited for other tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior3/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the tool scrapes a web page and returns a dictionary, which is useful, but it omits caveats about network dependency, page size limits, site accessibility, or what the vague 'additional info' actually contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness4/5Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-line summary, brief explanation, Args section, and Returns section. However, the first two sentences are somewhat redundant ('Count occurrences...' and 'This tool scrapes... counts...'), which slightly reduces tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness4/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with three parameters, the description covers inputs, behavior, and return type adequately. It does not specify the exact structure of the returned dictionary or note edge cases such as unavailable pages, but given the tool's simplicity, the remaining gaps are relatively minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters5/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for the schema's bare types. The Args section clearly explains all three parameters, including the url to analyze, the word to count, and the case_insensitive behavior with its default. It fully adds meaning beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource combination ('Count occurrences of a specific word on a web page') and clearly distinguishes the tool from siblings: scrape_web (raw content retrieval) and search_docs (searching documentation). The purpose is immediately obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose makes the tool's usage context clear: use it when you need word frequency from a web page. However, it does not explicitly contrast with alternatives or state exclusions for when another tool would be better, so it stops short of full usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently states that the tool downloads a zip, indexes markdown files, caches the index, and returns document previews. It does not mention network failure modes or cache invalidation, but the core side effects are clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness5/5Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a summary, Args, Returns, and Examples sections. Every section contributes useful information, and the core purpose is front-loaded in the first sentence without unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness5/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and moderate complexity, the description covers the full workflow, parameter semantics, defaults, return shape, and practical examples. The presence of an output schema covers the return format details, so nothing critical is missing for an agent to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters5/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Given 0% schema description coverage, the description fully compensates by explaining each parameter: query as a search string, zip_url with a default and required URL format, and num_results as a maximum count. This adds substantial meaning beyond the raw schema types and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches documentation from a GitHub repository zip file, with a specific verb, resource, and workflow (download, index, return relevant documents). It implicitly differentiates from siblings like scrape_web and count_word_occurrences by focusing specifically on GitHub-hosted documentation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines3/5Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through examples and the default zip_url, showing when to call the tool directly. However, there is no explicit guidance about when to prefer this tool over alternatives or any exclusions for edge cases like searching non-markdown documentation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
GitHub Badge
Glama performs regular codebase and documentation scans to:
- Confirm that the MCP server is working as expected.
- Confirm that there are no obvious security issues.
- Evaluate tool definition quality.
Our badge communicates server capabilities, safety, and installation instructions.
Card Badge
Copy to your README.md:
Score Badge
Copy to your README.md:
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/wanyingng/docs-search-engine'
If you have feedback or need assistance with the MCP directory API, please join our Discord server