Skip to main content
Glama
akhileshvj

Research Server

by akhileshvj

MCPChatbotForPapers

An MCP-based paper search chatbot that connects a Gemini client to local MCP tools for arXiv search and paper metadata lookup.

What's Included

  • mcp_chatbot.py: interactive chatbot that connects to configured MCP servers and routes tool calls through Gemini

  • research_server.py: MCP server that searches arXiv and stores paper metadata locally

  • papers/: generated cache of paper search results, grouped by topic

  • server_config.json: MCP server launch configuration used by the chatbot

Related MCP server: arxiv-mcp-server

Requirements

  • Python 3.14 or newer

  • uv installed locally

  • A Google Gemini API key

Quick Start

  1. Clone the repo and enter the project directory.

  2. Create a local .env file in the project root and add your Google API key.

  3. Install dependencies with uv sync.

  4. Start the research server.

  5. Start the chatbot in a second terminal and ask a question.

git clone git@github.com:akhileshvj/MCPChatbotForPapers.git
cd MCPChatbotForPapers
uv sync
uv run python research_server.py
uv run python mcp_chatbot.py

Initialize the Project

Clone the repository and move into it:

git clone git@github.com:akhileshvj/MCPChatbotForPapers.git
cd MCPChatbotForPapers

Create and use the virtual environment managed by uv:

uv sync

If you prefer to install from the pinned requirements file instead of pyproject.toml, use:

uv pip install -r requirements.txt

Configure Environment Variables

Create a .env file in the project root and add your Gemini API key:

GOOGLE_API_KEY=your_google_genai_api_key

Keep this file local. It is not meant to be pushed to GitHub.

The chatbot loads environment variables with python-dotenv.

Install Dependencies

If you are starting from a clean environment, install the project dependencies with:

uv sync

That will install the packages listed in pyproject.toml, including:

  • google-genai

  • mcp[cli]

  • python-dotenv

  • arxiv

  • fastapi

  • uvicorn

Run the MCP Research Server

The research server exposes the paper search tools over MCP stdio:

uv run python research_server.py

Run the Chatbot

Start the interactive chatbot in a second terminal:

uv run python mcp_chatbot.py

The chatbot reads server_config.json, launches the configured MCP servers, and then waits for queries at the prompt.

Example Usage

Inside the chatbot, try prompts like:

Search papers about diffusion models
Find recent papers on quantum computing
Look up paper details for a saved paper ID

How It Works

  1. mcp_chatbot.py connects to the MCP servers listed in server_config.json.

  2. Gemini receives the available tool schemas.

  3. When Gemini requests a tool call, the chatbot routes it to the correct MCP server.

  4. research_server.py searches arXiv and stores results under papers/<topic>/papers_info.json.

Notes

  • papers/ is populated automatically when you run searches.

  • If you change server commands in server_config.json, restart the chatbot so it reloads the config.

  • The project currently uses local stdio-based MCP servers, so each server process must be runnable from the repository root.

Troubleshooting

  • If the chatbot cannot connect to Gemini, check that GOOGLE_API_KEY is set in .env.

  • If research_server.py fails to start, make sure arxiv and mcp are installed in the active environment.

  • If you see stale results, delete the relevant folder under papers/ and run the search again.

Available Tools

2 tools
extract_infoA

Search for stored local information about a specific paper by its ID.

Args: paper_id: The unique short ID of the paper.

ParametersJSON Schema
NameRequiredDescriptionDefault
paper_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It states the tool 'searches stored local information' which implies a read-only operation, but does not detail side effects, limitations, or what exactly 'stored local information' means. It adds modest context beyond the schema, but is not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences total, with the first stating the purpose and the second describing the parameter. No unnecessary words, and the most important information is front-loaded. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter, an output schema, and a sibling tool, the description is reasonably complete. It explains what the tool does and the input, but does not mention any prerequisites (e.g., that the paper must have been previously stored locally) or what happens if the ID is not found. The output schema likely covers return values, so the description does not need to. A 4 is appropriate for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the description adds meaningful context: it specifies that 'paper_id' is 'The unique short ID of the paper.' This clarifies the format and uniqueness, which goes beyond the schema's title and type. This compensates well for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool searches for stored local information about a specific paper by its ID. It uses a specific verb ('Search') and resource ('information about a paper'), and the sibling tool 'search_papers' likely searches across papers, implying a distinction. However, it does not explicitly differentiate from the sibling, so a 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any guidance on when to use this tool versus alternatives, nor does it mention prerequisites or scenarios where it should not be used. It only implicitly suggests use when a paper ID is available, but lacks explicit usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_papersC

Search for papers on arXiv based on a topic and store their information.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYes
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description mentions 'store their information', indicating a side effect of persisting data. However, it does not disclose details such as where storage occurs, whether results are returned, authentication needs, rate limits, or any other behavioral traits. With no annotations provided, the description carries full burden and falls short.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core action. It is efficient but could benefit from additional structure (e.g., listing parameters or side effects).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two parameters, no annotations, one sibling tool, and an existing output schema (not provided), the description is incomplete. It lacks details on parameter usage (especially max_results), return value (despite output schema existing), side effects of storage, and differentiation from sibling. The description does not adequately compensate for the schema's lack of description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must compensate. It specifies the 'topic' parameter implicitly ('based on a topic') but does not explain the 'max_results' parameter or any constraints/format for either parameter. The description adds little meaning beyond the schema's bare parameter names and types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Search for papers on arXiv based on a topic and store their information', which specifies the verb (search), resource (papers on arXiv), and action (store). However, it does not explicitly distinguish from the sibling tool 'extract_info', though the verb 'search' vs 'extract' provides implicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'extract_info', nor does it mention prerequisites, limitations, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedextract_info
    • First observedsearch_papers

TDQS

B3.2/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: one retrieves stored local information by paper ID, the other searches arXiv for papers by topic. No overlap or ambiguity.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern: 'extract_info' and 'search_papers'. No mixed conventions.

Tool Count3/5

With only 2 tools, the server feels under-scoped for a 'Research Server'. More tools like list, delete, or update would be expected, but the count is not extreme.

Completeness3/5

The surface covers search and retrieval by ID, but lacks essential operations like listing all stored papers, deleting, or updating entries, leaving notable gaps.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    Not graded
    maintenance
    Enables AI assistants to search and retrieve academic papers from arXiv through MCP tools, supporting search by various criteria, detailed paper information, category browsing, and PDF content extraction.
    4
    129
    2
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables searching arXiv, fetching metadata, reading papers as section-aware Markdown, listing recent papers, and downloading PDFs via five MCP tools.
    23
    2
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/akhileshvj/MCPChatbotForPapers'

If you have feedback or need assistance with the MCP directory API, please join our Discord server