Skip to main content
Glama
Phicks-debug

linkedin-web-scrapper-mcp-server

by Phicks-debug

LinkedIn Web Scraper MCP Server

A Model Context Protocol (MCP) server that provides LinkedIn web scraping capabilities as tools for AI assistants. This server uses Playwright to automate LinkedIn people search and extract profile information, exposing these capabilities through the MCP protocol.

Features

  • MCP Tool Integration: Exposes LinkedIn scraping as MCP tools for AI assistants

  • People Search: Search LinkedIn profiles using keywords, location, and network filters

  • Profile Extraction: Extract profile names, URLs, and headlines from search results

  • Session Management: Automatic LinkedIn login with cookie persistence

  • Adaptive Selectors: Handles LinkedIn UI changes with multiple CSS selector strategies

  • Network Filtering: Filter by connection degree (1st, 2nd, 3rd+ connections)

  • Location Support: Filter by location using LinkedIn's geoUrn codes or location strings

Related MCP server: LinkedIn MCP Server

Installation

  1. Clone the repository:

git clone https://github.com/Phicks-debug/linkedin-web-scrapper.git
cd linkedin-web-scrapper-mcp-server
  1. Install dependencies:

npm install
  1. Install Playwright browsers:

npx playwright install
  1. Configure your LinkedIn credentials:

cp config.example.json config.json

Then edit config.json with your LinkedIn credentials:

{
  "linkedin": {
    "email": "your-linkedin-email@email.com",
    "password": "your-linkedin-password"
  },
  "browser": {
    "headless": false,
    "slowMo": 1000,
    "cookiesPath": "./cookies.json"
  }
}
  1. Build the server:

npm run build

Usage

As an MCP Server

This server is designed to be used with MCP-compatible AI assistants. The server exposes LinkedIn scraping functionality through the MCP protocol.

Starting the MCP Server

# Start the server (connects via stdio)
node dist/index.js

# For development with auto-rebuild
npm run watch

Using MCP Inspector (Development)

Test the server using the MCP Inspector:

npm run inspector

Available MCP Tools

search-linkedin-people

Search for LinkedIn profiles using web scraping.

Input Schema:

{
  "keywords": "software engineer", // Required: Keywords to search for
  "location": "105646813",        // Optional: Location filter (geoUrn or location string)
  "network": "F"                  // Optional: Network degree filter
}

Network Filter Options:

  • "F" - 1st degree connections only

  • "S" - 2nd degree connections

  • "O" - 3rd+ degree connections (out of network)

Location Examples:

  • "105646813" - Spain (using LinkedIn geoUrn)

  • "San Francisco" - Location string

  • Default: "104195383" if not specified

Response Format:

{
  "success": true,
  "count": 10,
  "profiles": [
    {
      "name": "John Doe",
      "profileUrl": "https://www.linkedin.com/in/johndoe",
      "headline": "Senior Software Engineer at Tech Company"
    }
  ],
  "filters": {
    "keywords": "software engineer",
    "location": "105646813",
    "network": "F"
  }
}

MCP Integration

Adding to Claude Desktop

Add this server to your Claude Desktop MCP configuration:

{
  "mcpServers": {
    "linkedin": {
      "command": "node",
      "args": ["/path/to/linkedin-web-scrapper-mcp-server/dist/index.js"],
      "cwd": "/path/to/linkedin-web-scrapper-mcp-server"
    }
  }
}

Using with Other MCP Clients

The server follows the standard MCP protocol and can be used with any MCP-compatible client by connecting to the stdio transport.

How It Works

  1. MCP Protocol: Exposes LinkedIn scraping as standardized MCP tools

  2. Browser Automation: Uses Playwright to control Chrome/Chromium browser

  3. Session Persistence: Saves LinkedIn session cookies to avoid repeated logins

  4. People Search: Navigates to LinkedIn people search with specified filters

  5. Profile Extraction: Extracts profile data using adaptive CSS selectors

  6. Structured Output: Returns JSON-formatted results via MCP protocol

Development

Scripts

Script

Description

npm run build

Compile TypeScript and make executable

npm run watch

Watch mode for development

npm run inspector

Launch MCP Inspector for testing

npm run dev

Build and run the server

Project Structure

├── index.ts              # Main MCP server implementation
├── config.json          # LinkedIn credentials and browser settings
├── cookies.json          # Saved session cookies (auto-generated)
├── package.json          # MCP server configuration
└── dist/                 # Compiled JavaScript output

Security & Privacy

  • Local Credentials: Your LinkedIn credentials are stored locally in config.json

  • Session Cookies: Saved locally in cookies.json for session persistence

  • No Data Transmission: No data is sent anywhere except to LinkedIn for scraping

  • Browser Automation: Uses a visible browser window to avoid detection

Technical Details

  • Protocol: Model Context Protocol (MCP) 0.6.0

  • Runtime: Node.js with TypeScript

  • Browser Engine: Playwright with Chromium

  • Transport: Standard I/O (stdio) for MCP communication

  • Target: LinkedIn People Search API

Error Handling

The server handles common scenarios:

  • Automatic LinkedIn login when session expires

  • LinkedIn security challenges (requires manual intervention)

  • UI changes through adaptive selectors

  • Network timeouts and connection issues

Limitations

  • LinkedIn Terms: Use responsibly and respect LinkedIn's terms of service

  • Rate Limiting: Avoid excessive requests to prevent detection

  • Manual Challenges: Security challenges require manual completion

  • UI Dependencies: May need updates if LinkedIn significantly changes their UI

License

MIT License - see LICENSE file for details.

Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Test with MCP Inspector

  5. Submit a pull request

For issues and feature requests, please use the GitHub issues page.

Available Tools

2 tools
scrape-linkedin-profileA

Scrape comprehensive data from a specific LinkedIn profile URL. Returns detailed profile information including experience, education, skills, licenses & certifications, and more.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileUrlYesThe LinkedIn profile URL to scrape (e.g., 'https://www.linkedin.com/in/username/')

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes what data is returned (experience, education, etc.) but does not disclose potential side effects, limitations (e.g., profile privacy, rate limits), or authentication requirements. For a scraping tool, such context is important but missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences clearly state the action and the output categories. No filler or unnecessary repetition. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no output schema, the description adequately explains the return values by listing key sections. However, it lacks any caveats about scraping LinkedIn (e.g., profile must be public, potential blocks) which would be valuable for completeness, but the tool is simple enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single parameter (profileUrl) 100%, including an example. The description adds no additional meaning beyond what the schema already provides. Baseline of 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific action ('Scrape comprehensive data from a specific LinkedIn profile URL') and the resource. It distinguishes from the sibling tool 'search-linkedin-people' by emphasizing 'specific URL' versus search functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: scrape when you have a specific LinkedIn profile URL. However, it does not explicitly state when not to use it or mention the alternative sibling tool (search-linkedin-people) for finding profiles. This is implied rather than explicitly guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search-linkedin-peopleB

Search for LinkedIn profiles using web scraping. Returns profile names, URLs, and headlines.

ParametersJSON Schema
NameRequiredDescriptionDefault
networkNoNetwork degree filter: 'F' = 1st degree connections, 'S' = 2nd degree connections, 'O' = 3rd+ degree connections
keywordsNoKeywords to search for in profiles (e.g., 'AI engineer', 'data scientist')
locationNoLocation filter - can be a location string (e.g., 'San Francisco') or LinkedIn geoUrn code (e.g., '105646813' for Spain). Default: '104195383'

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions 'web scraping' and the return fields, but lacks important context such as whether authentication is required, rate limits, pagination, or potential failures. The absence of such details leaves the agent with significant uncertainty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is immediately meaningful and front-loaded. Every word contributes to the understanding of the tool's function and output, with no waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with no output schema and no annotations, the description is too brief. It omits critical context such as the expected number of results, whether results are paginated, and any limitations of web scraping. The information provided is not enough for an agent to confidently rely on this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage of all three parameters with descriptions. The tool description adds no additional parameter semantics beyond what the schema already states, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search for LinkedIn profiles using web scraping.' The verb 'search' and resource 'LinkedIn profiles' are specific, and the method 'web scraping' sets it apart from the sibling tool 'scrape-linkedin-profile', which implies a different action on a specific profile.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling 'scrape-linkedin-profile'. While the name implies searching versus scraping, there is no explicit mention of alternatives, prerequisites, or cases where this tool should be avoided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv1.0.0
    • First observedscrape-linkedin-profile
    • First observedsearch-linkedin-people

TDQS

A3.6/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: one discovers profiles via search, the other extracts detailed data from a specific profile URL. There is no overlap in functionality.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern using 'search-linkedin-people' and 'scrape-linkedin-profile', maintaining the same structure and style.

Tool Count3/5

With only two tools, the server is minimal but each serves a necessary step in the LinkedIn people-scraping workflow. The count feels slightly thin, yet it aligns with the focused purpose of the server.

Completeness4/5

The tools form a complete flow: search for profiles and then scrape any selected profile. Minor gaps exist (e.g., no company or job search), but for a people-focused scraper, the core lifecycle is covered.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Phicks-debug/linkedin-web-scrapper-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server