Web Content MCP Server
The Web Content MCP Server fetches, processes, and provides web content as context for LLMs, leveraging Cloudflare Browser Rendering.
Fetch Page: Fetches and processes a web page, with options to limit content length and include a screenshot.
Search Documentation: Searches Cloudflare documentation and returns relevant results.
Extract Structured Content: Extracts specific content from web pages using CSS selectors.
Summarize Content: Summarizes web content for more concise LLM context.
Uses Cloudflare Browser Rendering to extract web content for LLM context through both REST API and Workers Binding API
Leverages Puppeteer through Cloudflare's implementation (@cloudflare/puppeteer) for browser automation and content extraction
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Content MCP Serverfetch the latest Cloudflare documentation on Browser Rendering"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Cloudflare Browser Rendering Experiments & MCP Server
This project demonstrates how to use Cloudflare Browser Rendering to extract web content for LLM context. It includes experiments with the REST API and Workers Binding API, as well as an MCP server implementation that can be used to provide web context to LLMs.
Project Structure
cloudflare-browser-rendering/
├── examples/ # Example implementations and utilities
│ ├── basic-worker-example.js # Basic Worker with Browser Rendering
│ ├── minimal-worker-example.js # Minimal implementation
│ ├── debugging-tools/ # Tools for debugging
│ │ └── debug-test.js # Debug test utility
│ └── testing/ # Testing utilities
│ └── content-test.js # Content testing utility
├── experiments/ # Educational experiments
│ ├── basic-rest-api/ # REST API tests
│ ├── puppeteer-binding/ # Workers Binding API tests
│ └── content-extraction/ # Content processing tests
├── src/ # MCP server source code
│ ├── index.ts # Main entry point
│ ├── server.ts # MCP server implementation
│ ├── browser-client.ts # Browser Rendering client
│ └── content-processor.ts # Content processing utilities
├── puppeteer-worker.js # Cloudflare Worker with Browser Rendering binding
├── test-puppeteer.js # Tests for the main implementation
├── wrangler.toml # Wrangler configuration for the Worker
├── cline_mcp_settings.json.example # Example MCP settings for Cline
├── .gitignore # Git ignore file
└── LICENSE # MIT LicenseRelated MCP server: MCP Web Tools Server
Prerequisites
Node.js (v16 or later)
A Cloudflare account with Browser Rendering enabled
TypeScript
Wrangler CLI (for deploying the Worker)
Installation
Clone the repository:
git clone https://github.com/yourusername/cloudflare-browser-rendering.git
cd cloudflare-browser-renderingInstall dependencies:
npm installCloudflare Worker Setup
Install the Cloudflare Puppeteer package:
npm install @cloudflare/puppeteerConfigure Wrangler:
# wrangler.toml
name = "browser-rendering-api"
main = "puppeteer-worker.js"
compatibility_date = "2023-10-30"
compatibility_flags = ["nodejs_compat"]
[browser]
binding = "browser"Deploy the Worker:
npx wrangler deployTest the Worker:
node test-puppeteer.jsRunning the Experiments
Basic REST API Experiment
This experiment demonstrates how to use the Cloudflare Browser Rendering REST API to fetch and process web content:
npm run experiment:restPuppeteer Binding API Experiment
This experiment demonstrates how to use the Cloudflare Browser Rendering Workers Binding API with Puppeteer for more advanced browser automation:
npm run experiment:puppeteerContent Extraction Experiment
This experiment demonstrates how to extract and process web content specifically for use as context in LLMs:
npm run experiment:contentMCP Server
The MCP server provides tools for fetching and processing web content using Cloudflare Browser Rendering for use as context in LLMs.
Building the MCP Server
npm run buildRunning the MCP Server
npm startOr, for development:
npm run devMCP Server Tools
The MCP server provides the following tools:
fetch_page- Fetches and processes a web page for LLM contextsearch_documentation- Searches Cloudflare documentation and returns relevant contentextract_structured_content- Extracts structured content from a web page using CSS selectorssummarize_content- Summarizes web content for more concise LLM context
Configuration
To use your Cloudflare Browser Rendering endpoint, set the BROWSER_RENDERING_API environment variable:
export BROWSER_RENDERING_API=https://YOUR_WORKER_URL_HEREReplace YOUR_WORKER_URL_HERE with the URL of your deployed Cloudflare Worker. You'll need to replace this placeholder in several files:
In test files:
test-puppeteer.js,examples/debugging-tools/debug-test.js,examples/testing/content-test.jsIn the MCP server configuration:
cline_mcp_settings.json.exampleIn the browser client:
src/browser-client.ts(as a fallback if the environment variable is not set)
Integrating with Cline
To integrate the MCP server with Cline, copy the cline_mcp_settings.json.example file to the appropriate location:
cp cline_mcp_settings.json.example ~/Library/Application\ Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.jsonOr add the configuration to your existing cline_mcp_settings.json file.
Key Learnings
Cloudflare Browser Rendering requires the
@cloudflare/puppeteerpackage to interact with the browser binding.The correct pattern for using the browser binding is:
import puppeteer from '@cloudflare/puppeteer'; // Then in your handler: const browser = await puppeteer.launch(env.browser); const page = await browser.newPage();When deploying a Worker that uses the Browser Rendering binding, you need to enable the
nodejs_compatcompatibility flag.Always close the browser after use to avoid resource leaks.
License
MIT
Available Tools
4 toolsextract_structured_contentC
Extracts structured content from a web page using CSS selectors
| Name | Required | Description | Default |
|---|---|---|---|
| selectors | Yes | CSS selectors to extract content | |
| url | Yes | URL to extract content from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but lacks details on traits like error handling (e.g., invalid URLs or selectors), performance (e.g., timeouts or rate limits), output format (e.g., JSON structure), or side effects (e.g., whether it makes network requests). This leaves significant gaps for an agent to understand how the tool behaves beyond its basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose ('Extracts structured content from a web page') and specifies the method ('using CSS selectors'). There is no wasted verbiage or redundant information, making it highly concise and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (involving web scraping with CSS selectors), lack of annotations, and no output schema, the description is incomplete. It doesn't cover behavioral aspects like error handling, output structure, or limitations (e.g., JavaScript-rendered content). While the schema documents parameters well, the overall context for safe and effective use is insufficient, especially for a tool that interacts with external resources.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for both parameters ('url' and 'selectors'). The description adds minimal value beyond the schema by mentioning 'CSS selectors', which aligns with the schema's description for 'selectors'. It doesn't provide additional context like selector syntax examples or URL validation rules. Given the high schema coverage, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('extracts') and target ('structured content from a web page'), and specifies the method ('using CSS selectors'). It distinguishes from siblings like 'fetch_page' (which likely retrieves raw HTML) and 'summarize_content' (which processes content). However, it doesn't explicitly contrast with 'search_documentation', leaving some ambiguity about when to choose one over the other.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a valid URL), exclusions (e.g., not for non-web content), or comparisons with sibling tools like 'fetch_page' for raw HTML or 'search_documentation' for query-based extraction. Usage is implied only by the tool's name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_pageC
Fetches and processes a web page for LLM context
| Name | Required | Description | Default |
|---|---|---|---|
| includeScreenshot | No | Whether to include a screenshot (base64 encoded) | |
| maxContentLength | No | Maximum content length to return | |
| url | Yes | URL to fetch |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'fetches and processes' but doesn't clarify aspects like rate limits, authentication needs, error handling, or what 'processes' entails (e.g., cleaning HTML, extracting text). This leaves significant gaps for a tool that interacts with external resources.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without any wasted words. It's appropriately sized for the tool's complexity, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (fetching and processing web pages), lack of annotations, and no output schema, the description is incomplete. It doesn't explain return values, error cases, or behavioral traits, leaving the agent with insufficient information for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, so the baseline is 3. The description adds no additional meaning beyond the schema, which already details parameters like 'url', 'includeScreenshot', and 'maxContentLength'. No extra syntax, format, or usage context is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('fetches and processes') and resource ('a web page'), and specifies the outcome ('for LLM context'). It doesn't explicitly differentiate from sibling tools like 'extract_structured_content' or 'summarize_content', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'extract_structured_content' or 'search_documentation'. It lacks context about prerequisites, exclusions, or comparative use cases, offering only a basic functional statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_documentationC
Searches Cloudflare documentation and returns relevant content
| Name | Required | Description | Default |
|---|---|---|---|
| maxResults | No | Maximum number of results to return | |
| query | Yes | Search query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions 'returns relevant content' but doesn't specify what 'relevant' means, whether results are ranked, if there's pagination, rate limits, authentication requirements, or what format/content is returned. This leaves significant behavioral aspects unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core functionality without unnecessary words. It's appropriately sized and front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no annotations and no output schema, the description is inadequate. It doesn't explain what constitutes 'relevant content', how results are structured, whether there are limitations or constraints, or how this differs from sibling tools that also retrieve content.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters clearly documented in the schema. The description doesn't add any meaningful parameter semantics beyond what the schema already provides, so it meets the baseline of 3 for high schema coverage without adding value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Searches') and target resource ('Cloudflare documentation'), providing a specific verb+resource combination. However, it doesn't differentiate from sibling tools like 'fetch_page' or 'extract_structured_content' which might also retrieve documentation content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives like 'fetch_page' for direct retrieval or 'extract_structured_content' for processing. The description only states what it does, not when it's appropriate or what distinguishes it from sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_contentC
Summarizes web content for more concise LLM context
| Name | Required | Description | Default |
|---|---|---|---|
| maxLength | No | Maximum length of the summary | |
| url | Yes | URL to summarize |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the tool 'summarizes web content' but fails to describe key behaviors such as how the summarization is performed (e.g., algorithm, quality), potential rate limits, error handling, or what the output looks like. This leaves significant gaps in understanding the tool's operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any redundant information. It is appropriately sized and front-loaded, making it easy to understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't address the tool's behavior, output format, or error conditions, which are crucial for a tool that processes web content. The high schema coverage helps with parameters, but overall context is insufficient for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents both parameters ('url' and 'maxLength') adequately. The description adds no additional meaning beyond what the schema provides, such as explaining the units for 'maxLength' or constraints on the 'url'. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('summarizes') and resource ('web content'), and it provides the intended outcome ('for more concise LLM context'). However, it doesn't explicitly differentiate from sibling tools like 'extract_structured_content' or 'fetch_page', which might also process web content in different ways.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose summarization over extraction or fetching, nor does it specify any prerequisites or exclusions for usage, leaving the agent to infer context from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v1.0.0- First observed
extract_structured_content - First observed
fetch_page - First observed
search_documentation - First observed
summarize_content
TDQS
The tools are mostly distinct in purpose, with extract_structured_content, fetch_page, and summarize_content each handling different aspects of web content processing. However, fetch_page and summarize_content could be slightly confused as both relate to processing web content for LLM context, though their specific focuses differ.
All tool names follow a consistent verb_noun pattern with snake_case, such as extract_structured_content and fetch_page. This uniformity makes the set predictable and easy to understand, with no deviations in naming conventions.
With 4 tools, the count is on the lower side for a web content server, which might feel thin for covering a broad domain like web content processing. However, it is reasonable for a focused set, though it could benefit from additional tools for more comprehensive coverage.
There are significant gaps in the tool surface for web content processing. For example, there are no tools for updating or deleting content, handling dynamic content like JavaScript, or managing multiple pages. The inclusion of search_documentation is also inconsistent with the general web content focus, creating a notable gap in core operations.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…
Zenrows MCP server — Fetch, Extract, Batch, and Browser Sessions for AI coding assistants
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Related MCP Servers
- AlicenseBqualityDmaintenanceA server that enables browser automation using Playwright, allowing interaction with web pages, capturing screenshots, and executing JavaScript in a browser environment through LLMs.1218,1221MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that allows LLMs to interact with web content through standardized tools, currently supporting web scraping functionality.1MIT
- AlicenseAqualityDmaintenanceThis MCP server provides tools for interacting with Cloudflare Browser Rendering, allowing you to fetch and process web content for use as context in LLMs directly from Cline or Claude Desktop.511MIT
- FlicenseNot gradedqualityDmaintenanceA server that enables AI systems to browse, retrieve content from, and interact with web pages through the Model Context Protocol.1-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/amotivv/cloudflare-browser-rendering'
If you have feedback or need assistance with the MCP directory API, please join our Discord server