Skip to main content
Glama
ethan-tsai-tsai

Thread Analyzer MCP Server


What is this?

An MCP server that lets your AI assistant scrape a Threads post URL, then query and analyze the replies — all through natural conversation.

Instead of:

1. Manually open browser
2. Scroll through hundreds of replies
3. Copy-paste into spreadsheet
4. Manually look for patterns

Just say:

"Analyze the replies on this Threads post: https://www.threads.com/@zuck/post/ABC123"

Your AI agent handles the rest.

Related MCP server: threads-mcp

Features

  • Network interception — Captures GraphQL API responses, not fragile CSS selectors that Meta randomizes

  • Anti-detection — Randomized scroll delays, stealth browser flags, custom User-Agent

  • 4 MCP tools — Scrape, list, search, and get statistics on replies

  • Works with any MCP client — Claude Code, Claude Desktop, Cursor, VS Code, Windsurf, and more

MCP Tools

Tool

Description

scrape_thread(url)

Scrape all replies from a public Threads post

get_all_replies()

Return all scraped replies with username and timestamp

search_replies(keyword)

Case-insensitive keyword search across replies

get_reply_stats()

Reply count, top commenters, avg length, time range

Quick Start

Prerequisites

  • Python 3.13+

  • uv package manager

Installation

git clone https://github.com/ethan-tsai-tsai/thread-analyzer.git
cd thread-analyzer
uv sync
uv run playwright install chromium

Configuration

Add the server to your MCP client config:

{
  "mcpServers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Add to your project's .mcp.json:

{
  "mcpServers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Or run: claude mcp add thread-analyzer -- uv run --directory /absolute/path/to/thread-analyzer python server.py

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Go to Cursor Settings > MCP > Add new MCP Server, then add:

{
  "mcpServers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Add to .vscode/mcp.json in your workspace:

{
  "servers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Add to ~/.codeium/windsurf/mcp_config.json:

{
  "mcpServers": {
    "thread-analyzer": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/thread-analyzer", "python", "server.py"]
    }
  }
}

Standalone CLI

You can also use the scraper directly without MCP:

uv run python scraper.py "https://www.threads.com/@user/post/XXXXX"

# Options
uv run python scraper.py "URL" --output custom.csv --max-scrolls 50

How It Works

┌─────────────┐     MCP (stdio)     ┌──────────────┐    Playwright    ┌─────────────┐
│  AI Client  │ ◄──────────────────► │  server.py   │ ◄──────────────► │  Threads.com │
│ (Claude,    │   scrape_thread()    │  (FastMCP)   │   GraphQL API   │  (Meta)      │
│  Cursor...) │   get_all_replies()  │              │   interception  │              │
│             │   search_replies()   │  replies.csv │                 │              │
│             │   get_reply_stats()  │              │                 │              │
└─────────────┘                      └──────────────┘                 └─────────────┘
  1. You give your AI assistant a Threads post URL

  2. AI calls scrape_thread(url) via MCP

  3. Server launches headless Chromium, navigates to the post

  4. Playwright intercepts GraphQL network responses containing reply data

  5. Server parses replies (username, text, timestamp), saves to CSV

  6. AI uses get_all_replies(), search_replies(), get_reply_stats() to analyze

Anti-Detection

Technique

Purpose

Custom User-Agent

Mimics real Chrome browser

navigator.webdriver removal

Hides automation flag

AutomationControlled disabled

Prevents Chromium detection

Randomized scroll delays (1.5-4.5s)

Avoids behavioral fingerprinting

Early stop on idle scrolls

Mimics natural browsing patterns

Limitations

  • Public posts only — Cannot access private or restricted posts

  • Meta's anti-bot measures — Meta may block headless browsers; if scraping fails, try the standalone CLI in non-headless mode

  • GraphQL schema changes — Meta periodically changes their API structure; the parser in scraper.py may need updating

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

License

MIT


Available Tools

4 tools
get_all_repliesA

Get all scraped Threads replies.

Returns a formatted list of all replies with username, timestamp and text. Call scrape_thread first if no data exists yet.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the return format (formatted list with username, timestamp, text) and the dependency on scrape_thread. However, it doesn't specify behavior on empty data or potential errors, which is a notable gap for a data retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and front-loaded with the primary action. Every sentence earns its place, covering purpose, output format, and necessary prerequisite without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple nature (0 params, output schema exists), the description is fairly complete. It covers the return fields and the dependency on scrape_thread. The only gap is lack of explicit error/empty-list behavior, but overall it's sufficient for a straightforward getter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds context about the output and prerequisite, which is helpful even though there are no parameters to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gets all scraped Threads replies, with a specific verb and resource. It distinguishes itself from sibling tools like search_replies by emphasizing 'all' replies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit prerequisite ('Call scrape_thread first if no data exists yet'), giving context on when to use the tool. It doesn't explicitly exclude alternatives or list when not to use it, but the 'all' wording implies a contrast with filtered searching.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reply_statsA

Get statistics about the scraped replies.

Returns reply count, unique users, most active commenters, and time range of replies.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the burden of disclosing side effects and operational context. It mentions 'scraped replies,' implying a read-only operation on stored data, and lists return values. However, it does not explicitly state that the tool is read-only, requires prior scraping, or has no side effects, leaving some ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded. The first sentence states the core purpose, and the second lists the output fields. No wasted words or redundant information, maintaining excellent structure for quick parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless tool with an output schema, the description is nearly complete. It lists the key return stats. However, it leaves minor context gaps, such as the prerequisite of prior scraping or the exact data scope, which could be clarified to make it fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to convey. As per the rubric, the baseline for 0-param tools is 4, and the description adds no unnecessary parameter detail since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool does: 'Get statistics about the scraped replies.' It then enumerates the specific stats returned (reply count, unique users, most active commenters, time range), making it clear and distinct from sibling tools that scrape threads, fetch all replies, or search replies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit. The description indicates it operates on 'scraped replies,' suggesting it is meant for post-scrape analysis, but it does not explicitly say when to prefer this over alternatives like get_all_replies or search_replies. No exclusions or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrape_threadA

Scrape replies from a public Threads post URL.

Opens a browser, navigates to the post, intercepts GraphQL API responses to extract replies, and saves them locally. This must be called before using get_all_replies or search_replies.

Args: url: Full URL of the Threads post (e.g. https://www.threads.com/@user/post/ABC123) max_scrolls: Maximum scroll iterations for loading more replies (default: 20)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_scrollsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool opens a browser, navigates, intercepts GraphQL responses, and saves data locally—behavioral traits beyond the schema. It also notes the 'public' constraint. However, it does not detail potential side effects like rate limits, file overwrites, or cleanup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: purpose, mechanism, prerequisite, then a simple argument list. Every sentence serves a function, and it front-loads the key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (browser automation, local save) and that an output schema exists, the description covers the prerequisite relationship and parameter details. It could be more complete by noting potential failure modes, but it is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It does: url is described as a 'Full URL' with an example, and max_scrolls as 'Maximum scroll iterations for loading more replies (default: 20).' This adds meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states its function: 'Scrape replies from a public Threads post URL.' It also details the mechanism (opens a browser, intercepts GraphQL API responses) and differentiates itself from siblings by noting it must precede get_all_replies/search_replies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a prerequisite condition: 'This must be called before using get_all_replies or search_replies,' which tells the agent when to use it relative to other tools. It also implies it is the initial scrape step and works only on public posts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_repliesA

Search Threads replies for a keyword (case-insensitive).

Args: keyword: The search term to filter replies by.

Returns matching replies with username and text.

ParametersJSON Schema
NameRequiredDescriptionDefault
keywordYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It discloses the case-insensitive behavior and the return fields (username and text), but does not explicitly state that it is read-only, mention any side effects, or clarify scope (e.g., all threads or a specific thread).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise, with the primary purpose stated in the first sentence and a simple args section. No redundant information is present; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter, and the description provides the essential information: what it does, the parameter meaning, and return structure. Given the expected output schema, it does not need to elaborate further. Minor gaps like scope or sorting are not critical for basic usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description compensates by explaining the 'keyword' parameter as 'the search term to filter replies by'. This adds meaning beyond the raw schema field name, though it could be more detailed (e.g., accepted formats or behaviors).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it searches Threads replies by a keyword, with a specific verb and resource. It distinguishes itself from siblings like 'get_all_replies' by focusing on keyword filtering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as 'get_all_replies' or 'get_reply_stats'. The description simply describes the function without contextualizing it among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv0.1.0
    • First observedget_all_replies
    • First observedget_reply_stats
    • First observedscrape_thread
    • First observedsearch_replies

TDQS

A4.3/5.0
Disambiguation5/5

Each tool has a distinct role: scrape_thread for fetching, get_all_replies for retrieving all, search_replies for filtering, get_reply_stats for analytics. No overlap or ambiguity.

Naming Consistency5/5

All tools follow a clear verb_noun snake_case pattern (scrape_, get_, search_), making the API predictable and easy to navigate.

Tool Count5/5

Four tools is well-scoped for a focused Threads analysis server, covering the full workflow without bloat or missing essentials.

Completeness5/5

The tool surface covers scraping, retrieval, search, and stats—everything needed for the stated purpose of analyzing Threads replies. No obvious gaps.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Manage Threads and Bluesky social media from AI assistants. Schedule posts, check analytics, and automate follow-up replies.
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for Threads comment research — analyze public sentiment, frustrations, and conversations on Threads via AI agents.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Lets coding agents fetch real social media and web data from platforms like TikTok, Instagram, YouTube, and more, directly inside editors like Cursor and VS Code.
    1
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ethan-tsai-tsai/thread-analyzer'

If you have feedback or need assistance with the MCP directory API, please join our Discord server