InfinityScrape MCP
Scans Codeforces for public user profiles and competitive programming activity as part of multi-platform OSINT reconnaissance.
Scans Dev.to for public developer profiles and posts as part of multi-platform OSINT reconnaissance.
Provides real-time web searches via DuckDuckGo, including text search and automatic scraping of top results.
Scans GitHub for public user, team, and repository information as part of multi-platform OSINT reconnaissance.
Scans GitLab for public user, group, and project information as part of multi-platform OSINT reconnaissance.
Scans Google Scholar for scholarly profiles and publications as part of multi-platform OSINT reconnaissance.
Scans Kaggle for public user profiles and dataset/competition activity as part of multi-platform OSINT reconnaissance.
Scans LeetCode for public user profiles and coding activity as part of multi-platform OSINT reconnaissance.
Scans Medium for public author profiles and articles as part of multi-platform OSINT reconnaissance.
Resolves addresses to GPS coordinates, street/postcode details, and administrative boundaries using OpenStreetMap.
Scans Reddit for public user profiles and activity as part of multi-platform OSINT reconnaissance.
Scans ResearchGate for public researcher profiles and publications as part of multi-platform OSINT reconnaissance.
Scans Substack for public author profiles and newsletters as part of multi-platform OSINT reconnaissance.
Extracts YouTube video, Shorts, and live transcripts with timestamps directly over HTTP, without requiring video downloads or local GPU models.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@InfinityScrape MCPScrape the content from https://example.com and summarize it, bypassing anti-bot checks."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ InfinityScrape MCP: World-Class Web Scraping, Dynamic SPA Rendering & 25-Tool OSINT Intelligence Suite
InfinityScrape MCP is a standalone, production-grade Model Context Protocol (MCP) server engineered to provide AI models (Open WebUI, Claude 3.7, DeepSeek-R1/V3, Antigravity AI, Cursor, LM Studio) with unlimited, high-speed, anti-bot resilient web scraping, dynamic SPA rendering, DuckDuckGo web search, Wayback Machine time-travel, instant YouTube transcription, and precision OSINT / GEOINT location intelligence.
๐ Table of Contents
Related MCP server: FineData MCP Server
๐ Why InfinityScrape MCP?
Standard web scrapers often fail on modern websites due to Cloudflare challenges, heavy client-side JavaScript rendering, intrusive cookie consent modals, and rate limits. InfinityScrape solves these problems out-of-the-box:
Dual-Engine Scraping Architecture:
Fast TLS Engine (
primp+httpx): Mimics real Chrome/Safari browser TLS/JA3 fingerprints and HTTP/2 headers to bypass Cloudflare and Akamai challenges in<100ms.Dynamic Headless Browser (
Playwright Chromium): Renders complex SPAs (React, Vue, Next.js, Angular), performs infinite scrolling, clicks elements, and executes custom JavaScript.
Network-Level Ad & Tracker Elimination:
Intercepts and aborts network calls to 35+ ad networks and tracking scripts (
doubleclick,criteo,outbrain,google-analytics) before they download, cutting page load time by ~300% and memory usage by 70%.Automatically detects and decomposes OneTrust, Cookiebot, and sticky overlay popups.
Zero-API Key Real-Time Web Search:
Native DuckDuckGo live search and Google Dork query generation to conduct up-to-date research without paying for Serper, Brave, or Bing search APIs.
Wayback Machine Time-Travel:
Query internet archive history for any URL across custom date ranges to track competitor pricing changes, deleted pages, and historical copy.
Zero-GPU Instant YouTube Transcriber:
Extracts complete video/shorts/live transcripts with timestamps (
[MM:SS]) in<300msdirectly via HTTP streams without downloading video or requiring local GPU Whisper models.
Deep Recursive Documentation Crawler:
Asynchronous Breadth-First-Search (BFS) crawler with domain locking and path prefix filtering to aggregate entire documentation trees into unified Markdown.
State-of-the-Art Public OSINT & GEOINT Reconnaissance:
Multi-Signal Confidence Scoring (0% - 100%): Evaluates Name + City + Street + PIN + Org + Role correlation to rank discovered dossiers.
25+ Global Platform Scanners: Scans GitHub, GitLab, StackOverflow, Kaggle, HuggingFace, LeetCode, Codeforces, Dev.to, Medium, Substack, Google Scholar, ResearchGate, Reddit, etc.
OpenStreetMap GEOINT: Resolves global addresses down to street/postcode level with GPS coordinates and administrative boundaries.
SQLite Persistent Caching Layer:
In-memory and SQLite-backed local cache for instant
0msresponses on repeat lookups with configurable TTL.
โก Competitive Comparison
Feature / Capability | Standard MCP Scrapers | Cloud Scraping APIs | InfinityScrape MCP |
Cost & API Keys | Free (Basic) | Paid ($20 - $200/mo) | 100% Free / Zero API Keys |
Cloudflare / Akamai TLS Bypass | โ Fails / 403 | โ Yes | โ
Built-in ( |
Dynamic SPAs & Infinite Scroll | โ Limited | โ Yes | โ
Built-in ( |
Real-Time Web Search & Dorking | โ No | โ ๏ธ Extra Cost | โ Built-in (DuckDuckGo & Dorks) |
Wayback Historical Snapshots | โ No | โ No | โ Built-in (Archive API) |
Network-Level Ad & Popup Stripping | โ No | โ ๏ธ Partial | โ Built-in (35+ domains) |
Zero-GPU YouTube Transcripts | โ No | โ No | โ Built-in (<300ms) |
Online PDF Page-by-Page Parser | โ No | โ ๏ธ Extra Cost | โ
Built-in ( |
Deep Documentation Crawler | โ No | โ ๏ธ Extra Cost | โ Built-in (Async BFS) |
25+ Platform OSINT & Geocoding | โ No | โ No | โ Built-in (0-100% Confidence) |
OpenAPI 3.1.0 REST Bridge (Port 8000) | โ No | โ ๏ธ Proprietary | โ Built-in (FastAPI /docs) |
๐๏ธ Architectural Overview
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ AI Clients: Open WebUI / Claude Desktop / Cursor โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ
โผ โผ
[OpenAPI Bridge (Port 8000)] [Stdio JSON-RPC 2.0 Server]
FastAPI /docs & /openapi.json (server.py)
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ โผ
[Fast TLS Engine] [Playwright Engine] [OSINT / GEOINT] [Search & Media]
โข primp JA3/TLS โข Stealth Chromium โข 25+ Platform โข DuckDuckGo Search
โข HTTP/2 Headers โข Ad/Tracker Blocker Scanners โข Wayback Snapshots
โข <100ms Execution โข Infinite Scroll โข OpenStreetMap โข YouTube (<300ms)
โข Auto-Dismiss CMPs โข Reverse Geocoding โข Remote PDF Parser
โข Match Confidence
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SQLite Caching Layer (0ms) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ Quick Start & 1-Click Installation
1. Automated Setup
# Clone the repository
git clone https://github.com/virajverse/infinity-scraper.git
cd infinity-scraper
# Create virtual environment & install
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
pip install -e .
playwright install chromium2. Launch FastAPI Bridge (Port 8000)
python openapi_bridge.pyInteractive Swagger Docs:
http://127.0.0.1:8000/docsOpenAPI 3.1.0 Schema:
http://127.0.0.1:8000/openapi.json
๐ AI Client Integration
1. Open WebUI (FastAPI Bridge on Port 8000)
Ensure the bridge is running (
python openapi_bridge.py).In Open WebUI, navigate to Workspace -> Tools -> Add Tool.
Import from URL:
http://127.0.0.1:8000/openapi.jsonor useinfinity_scraper_suite.All 25 tools are instantly accessible to your agents!
2. Antigravity AI / Claude Desktop (Native Stdio)
Add to your mcp_config.json:
{
"mcpServers": {
"infinity-scraper": {
"command": "python",
"args": ["-m", "infinity_scraper.server"],
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}๐ ๏ธ Complete 25-Tool Reference Catalog
1. Real-Time Web Search & Archive OSINT (4 Tools)
Tool | Description |
| Real-time web search without API keys. Returns ranked URLs, titles, and snippets. |
| Generates advanced Google Dork strings ( |
| Fetches historical snapshots, archive timestamps, and past versions of any URL. |
| Audits HTTP security headers (CSP, HSTS, X-Frame-Options, CORS). |
2. Anti-Bot Web Scraping & Documentation Crawlers (9 Tools)
Tool | Description |
| Production-grade scraper with auto-switching (Fast TLS -> Playwright Chromium fallback). |
| Extracts clean, readable Markdown from any webpage, stripping navigation and ads. |
| Extracts HTML tables and converts them into structured Markdown / JSON datasets. |
| Returns complete raw HTML of a target page for custom DOM parsing. |
| Extracts OpenGraph, Twitter cards, JSON-LD schemas, and meta tags. |
| Extracts all visible text alongside outbound internal/external hyperlinks. |
| Recursive BFS documentation crawler with depth limits and domain locking. |
| Hybrid search-and-crawl engine to locate and summarize specific docs topics. |
| Fetches and compares two URLs, highlighting content diffs and added/removed text. |
3. Media, Video & Document Parsers (2 Tools)
Tool | Description |
| Zero-GPU, sub-300ms transcript extraction from YouTube videos, shorts, and live streams. |
| Streams and parses remote online PDF files page-by-page into Markdown text. |
4. Deep Public OSINT & Entity Reconnaissance (7 Tools)
Tool | Description |
| Scans 10+ developer platforms (GitHub, GitLab, StackOverflow, Kaggle, LeetCode). |
| Scans academic databases (Google Scholar, ResearchGate, arXiv, ORCID). |
| Scans public social networks and communities (Reddit, Dev.to, Medium, Substack). |
| Cross-validates emails, phone numbers, and social handles with confidence scoring. |
| Maps organizational structures, key leadership, and public company roles. |
| Aggregates multi-source OSINT dossiers on persons or organizations. |
| Converts free-form physical addresses into validated postal and geocoded records. |
5. Precision GEOINT & Image Intelligence (3 Tools)
Tool | Description |
| Forward geocoding via OpenStreetMap/Nominatim down to street and postal code. |
| Converts GPS coordinates (latitude, longitude) into full postal addresses. |
| Generates reverse image search query URLs for Google, Bing, Yandex, and TinEye. |
๐ป Command-Line Interface (CLI)
InfinityScrape provides a built-in CLI for quick terminal testing:
# Scrape a webpage into Markdown
infinity-scrape scrape "https://news.ycombinator.com" --format markdown
# Search DuckDuckGo from the terminal
infinity-scrape search "Generative Engine Optimization 2026" --limit 5
# Extract YouTube Transcript
infinity-scrape youtube "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
# OSINT Persona Lookup
infinity-scrape osint --name "Linus Torvalds" --platforms github,gitlab๐งช Running Automated Tests
# Run unit and integration tests
pytest tests/ -v๐ License & Authors
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
No tool schema history has been recorded yet.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Direct access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
- GoroOAuthai.usegoro
62 real-world tools for agents: search, scraping, social, enrichment, image, video, voice.
Live web access for agents: scrape, SERP search, crawl/map, 74 collectors, datasets, proxies.
Give your agent live data from Twitter, Reddit, the web and GitHub. No API keys, no scraping stack.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to perform undetectable browser automation that bypasses Cloudflare, antibots, and social media blocks. Provides 105 tools for element extraction, network debugging, and real-world web scraping with a 98.7% success rate on protected sites.1,883MIT
- AlicenseAqualityBmaintenanceEnables AI agents to scrape any website by providing tools for JavaScript rendering, antibot bypass, and automatic captcha solving. It supports synchronous, asynchronous, and batch scraping operations with built-in proxy rotation.5207MIT

ScrapeLab MCPofficial
AlicenseNot gradedqualityDmaintenanceEnables undetectable web scraping and browser automation for AI agents with 84 tools including stealth navigation, element extraction, network interception, and auto cookie consent dismissal. Bypasses anti-bot systems like Cloudflare and DataDome while providing LLM-ready markdown output and full Chrome DevTools Protocol access.MIT
HasData MCP Serverofficial
AlicenseAqualityAmaintenanceDirect access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.8425MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/virajverse/infinity-scraper'
If you have feedback or need assistance with the MCP directory API, please join our Discord server