Web Scout MCP Server
This server enables privacy-focused web searching via DuckDuckGo and clean content extraction from web pages.
Search the web: Submit any query and receive structured results (titles, URLs, snippets) with up to 25 results per search (default: 10).
Extract content from a single URL: Fetch clean, readable text from any webpage by stripping scripts, styles, and navigation elements.
Extract content from multiple URLs simultaneously: Pass an array of URLs to fetch and extract content in parallel.
Built-in reliability features: Includes rate limiting to avoid API blocks, memory optimization to prevent crashes, and robust error handling for consistent operation.
Enables privacy-focused web search queries through DuckDuckGo's search engine, returning structured search results with titles, URLs, and snippets, along with the ability to extract clean, readable content from resulting web pages.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Scout MCP Serversearch for recent breakthroughs in quantum computing and extract content from the top 3 results"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
✨ Features
🔍 DuckDuckGo Search: Fast and privacy-focused web search capability
📄 Content Extraction: Clean, readable text extraction from web pages
🚀 Parallel Processing: Support for extracting content from multiple URLs simultaneously
💾 Memory Optimization: Smart memory management to prevent application crashes
⏱️ Rate Limiting: Intelligent request throttling to avoid API blocks
🛡️ Error Handling: Robust error handling for reliable operation
Related MCP server: Web Search MCP Server
📦 Installation
Installing via Smithery
To install Web Scout for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @pinkpixel-dev/web-scout-mcp --client claudeGlobal Installation
npm install -g @pinkpixel/web-scout-mcpLocal Installation
npm install @pinkpixel/web-scout-mcp🚀 Usage
Command Line
After installing globally, run:
web-scout-mcpWith MCP Clients
Add this to your MCP client's config.json (Claude Desktop, Cursor, etc.):
{
"mcpServers": {
"web-scout": {
"command": "npx",
"args": [
"-y",
"@pinkpixel/web-scout-mcp@latest"
]
}
}
}Environment Variables
Set the WEB_SCOUT_DISABLE_AUTOSTART=1 environment variable when embedding the package and calling createServer() yourself. By default running the published entrypoint (for example node dist/index.js or npx @pinkpixel/web-scout-mcp) automatically bootstraps the stdio transport.
🧰 Tools
The server provides the following MCP tools:
🔍 DuckDuckGoWebSearch
Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings.
Input:
query(string): The search query stringmaxResults(number, optional): Maximum number of results to return (default: 10)
Example:
{
"query": "latest advancements in AI",
"maxResults": 5
}Output: A formatted list of search results with titles, URLs, and snippets.
📄 UrlContentExtractor
Fetches and extracts clean, readable content from web pages by removing unnecessary elements like scripts, styles, and navigation.
Input:
url: Either a single URL string or an array of URL strings
Example (single URL):
{
"url": "https://example.com/article"
}Example (multiple URLs):
{
"url": [
"https://example.com/article1",
"https://example.com/article2"
]
}Output: Extracted text content from the specified URL(s).
🛠️ Development
# Clone the repository
git clone https://github.com/pinkpixel-dev/web-scout-mcp.git
cd web-scout-mcp
# Install dependencies
npm install
# Build
npm run build
# Run
npm start📚 Documentation
For more detailed information about the project, check out these resources:
OVERVIEW.md - Technical overview and architecture
CONTRIBUTING.md - Guidelines for contributors
CHANGELOG.md - Version history and changes
📋 Requirements
Node.js >= 18.0.0
npm or yarn
📄 License
This project is licensed under the Apache 2.0 License.
Available Tools
2 toolsDuckDuckGoWebSearchCInspect
Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings. Input the keywords, question, or topic you want to search for using DuckDuckGo as your query. Input the maximum number of search entries you'd like to receive using maxResults - defaults to 10 if not provided.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query string | |
| maxResults | No | Maximum number of results to return (default: 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool 'returns a well-structured list of findings' and defaults maxResults to 10, but lacks details on rate limits, authentication needs, error handling, or what 'well-structured' entails. For a search tool with no annotation coverage, this leaves significant gaps in understanding its operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized with two sentences that efficiently cover the tool's function and parameters. It is front-loaded with the core purpose, though the second sentence could be slightly more streamlined by avoiding repetition of 'Input'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, no output schema, no annotations), the description is adequate but incomplete. It covers the basic purpose and parameters but lacks details on output format, error cases, and behavioral traits. Without annotations or an output schema, more context on what 'well-structured list' means would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters (query and maxResults) fully documented in the schema. The description adds minimal value beyond the schema: it reiterates that query is for 'keywords, question, or topic' and notes the default for maxResults, which is already in the schema. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings.' This specifies the verb ('initiates a web search'), resource ('web search query'), and engine ('DuckDuckGo'), distinguishing it from the sibling tool UrlContentExtractor. However, it doesn't explicitly contrast with the sibling beyond mentioning the engine.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like UrlContentExtractor. It mentions the query input and maxResults default but offers no context about appropriate use cases, prerequisites, or exclusions. Usage is implied through parameter descriptions but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
UrlContentExtractorCInspect
Fetches and extracts content from a given webpage URL. Input the URL of the webpage you want to extract content from as a string using the url parameter. You can also input an array of URLs to fetch content from multiple pages at once.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL or list of URLs to fetch |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool fetches and extracts content but lacks details on potential issues like rate limits, authentication needs, error handling, or what 'extracts content' entails (e.g., text, HTML, metadata). This leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, with the core purpose stated first. Both sentences are relevant, but the second sentence could be slightly more concise by combining the single and multiple URL explanations without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete for a tool that performs web content extraction. It doesn't explain what 'extracts content' means in terms of output format, potential limitations (e.g., JavaScript-rendered content), or error scenarios, leaving the agent with insufficient context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already fully documents the 'url' parameter as a string or array of URIs. The description adds minimal value by restating this in plain language without providing additional context, such as URL format constraints or performance implications of array inputs, aligning with the baseline score for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('fetches and extracts content') and resource ('from a given webpage URL'), making it easy to understand what it does. However, it doesn't explicitly differentiate from its sibling tool DuckDuckGoWebSearch, which likely serves a different search-oriented purpose rather than direct content extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as DuckDuckGoWebSearch. It mentions the ability to handle single or multiple URLs but doesn't clarify scenarios where one might prefer this over other tools or when it's inappropriate to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
- First observed
DuckDuckGoWebSearch - First observed
UrlContentExtractor
TDQS
The two tools have completely distinct purposes: DuckDuckGoWebSearch performs web searches to find URLs, while UrlContentExtractor extracts content from specific URLs. There is no overlap in functionality or ambiguity about when to use each tool.
Both tools use descriptive, multi-word names that clearly indicate their function. While not following a strict verb_noun pattern, they maintain readability and consistency in style. The minor deviation from perfect pattern consistency prevents a score of 5.
With only 2 tools for a web search and content extraction server, the surface feels thin and incomplete. A typical web search server would benefit from additional tools like advanced search filters, result pagination, or content analysis utilities to provide more comprehensive coverage.
While the basic search-to-extract workflow is covered, there are significant gaps in the web search domain. Missing operations include search result filtering, handling pagination, saving/search history, content summarization, or image/video search capabilities that would be expected in a complete web search toolkit.
Maintenance
Related MCP Connectors
Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Turn any URL into clean Markdown and structured data. Scrape, crawl, search and extract.
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables web searching through DuckDuckGo and fetching content from webpages. Provides search capabilities with configurable result limits and webpage content extraction for AI assistants.-
- FlicenseNot gradedqualityCmaintenanceEnables web search across multiple search engines (DuckDuckGo, Bing, Startpage) with parallel execution and result deduplication. Also provides web page content extraction capabilities.2-
- AlicenseBqualityDmaintenanceEnables web search through DuckDuckGo and webpage content fetching with intelligent text extraction. Features built-in rate limiting and LLM-optimized result formatting for seamless integration with language models.2MIT
- AlicenseAqualityBmaintenanceEnables web search without API keys using DuckDuckGo and Bing search engines, and retrieves webpage content. Supports multiple search engines simultaneously with privacy protection and asynchronous processing.29MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/pinkpixel-dev/web-scout-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server