Research Server
Searches arXiv for papers and stores metadata locally, enabling queries for recent papers on specific topics and retrieval of paper details.
Connects to Google Gemini as the client, routing tool calls from Gemini to MCP servers for paper search and lookup.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Research Serversearch for papers on reinforcement learning"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCPChatbotForPapers
An MCP-based paper search chatbot that connects a Gemini client to local MCP tools for arXiv search and paper metadata lookup.
What's Included
mcp_chatbot.py: interactive chatbot that connects to configured MCP servers and routes tool calls through Geminiresearch_server.py: MCP server that searches arXiv and stores paper metadata locallypapers/: generated cache of paper search results, grouped by topicserver_config.json: MCP server launch configuration used by the chatbot
Related MCP server: arxiv-mcp-server
Requirements
Python 3.14 or newer
uvinstalled locallyA Google Gemini API key
Quick Start
Clone the repo and enter the project directory.
Create a local
.envfile in the project root and add your Google API key.Install dependencies with
uv sync.Start the research server.
Start the chatbot in a second terminal and ask a question.
git clone git@github.com:akhileshvj/MCPChatbotForPapers.git
cd MCPChatbotForPapers
uv sync
uv run python research_server.py
uv run python mcp_chatbot.pyInitialize the Project
Clone the repository and move into it:
git clone git@github.com:akhileshvj/MCPChatbotForPapers.git
cd MCPChatbotForPapersCreate and use the virtual environment managed by uv:
uv syncIf you prefer to install from the pinned requirements file instead of pyproject.toml, use:
uv pip install -r requirements.txtConfigure Environment Variables
Create a .env file in the project root and add your Gemini API key:
GOOGLE_API_KEY=your_google_genai_api_keyKeep this file local. It is not meant to be pushed to GitHub.
The chatbot loads environment variables with python-dotenv.
Install Dependencies
If you are starting from a clean environment, install the project dependencies with:
uv syncThat will install the packages listed in pyproject.toml, including:
google-genaimcp[cli]python-dotenvarxivfastapiuvicorn
Run the MCP Research Server
The research server exposes the paper search tools over MCP stdio:
uv run python research_server.pyRun the Chatbot
Start the interactive chatbot in a second terminal:
uv run python mcp_chatbot.pyThe chatbot reads server_config.json, launches the configured MCP servers, and then waits for queries at the prompt.
Example Usage
Inside the chatbot, try prompts like:
Search papers about diffusion models
Find recent papers on quantum computing
Look up paper details for a saved paper IDHow It Works
mcp_chatbot.pyconnects to the MCP servers listed inserver_config.json.Gemini receives the available tool schemas.
When Gemini requests a tool call, the chatbot routes it to the correct MCP server.
research_server.pysearches arXiv and stores results underpapers/<topic>/papers_info.json.
Notes
papers/is populated automatically when you run searches.If you change server commands in
server_config.json, restart the chatbot so it reloads the config.The project currently uses local stdio-based MCP servers, so each server process must be runnable from the repository root.
Troubleshooting
If the chatbot cannot connect to Gemini, check that
GOOGLE_API_KEYis set in.env.If
research_server.pyfails to start, make surearxivandmcpare installed in the active environment.If you see stale results, delete the relevant folder under
papers/and run the search again.
Available Tools
2 toolsextract_infoA
Search for stored local information about a specific paper by its ID.
Args: paper_id: The unique short ID of the paper.
| Name | Required | Description | Default |
|---|---|---|---|
| paper_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility for behavioral disclosure. It states the tool 'searches stored local information' which implies a read-only operation, but does not detail side effects, limitations, or what exactly 'stored local information' means. It adds modest context beyond the schema, but is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences total, with the first stating the purpose and the second describing the parameter. No unnecessary words, and the most important information is front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one parameter, an output schema, and a sibling tool, the description is reasonably complete. It explains what the tool does and the input, but does not mention any prerequisites (e.g., that the paper must have been previously stored locally) or what happens if the ID is not found. The output schema likely covers return values, so the description does not need to. A 4 is appropriate for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description adds meaningful context: it specifies that 'paper_id' is 'The unique short ID of the paper.' This clarifies the format and uniqueness, which goes beyond the schema's title and type. This compensates well for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches for stored local information about a specific paper by its ID. It uses a specific verb ('Search') and resource ('information about a paper'), and the sibling tool 'search_papers' likely searches across papers, implying a distinction. However, it does not explicitly differentiate from the sibling, so a 4 is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives, nor does it mention prerequisites or scenarios where it should not be used. It only implicitly suggests use when a paper ID is available, but lacks explicit usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersC
Search for papers on arXiv based on a topic and store their information.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions 'store their information', indicating a side effect of persisting data. However, it does not disclose details such as where storage occurs, whether results are returned, authentication needs, rate limits, or any other behavioral traits. With no annotations provided, the description carries full burden and falls short.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the core action. It is efficient but could benefit from additional structure (e.g., listing parameters or side effects).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given two parameters, no annotations, one sibling tool, and an existing output schema (not provided), the description is incomplete. It lacks details on parameter usage (especially max_results), return value (despite output schema existing), side effects of storage, and differentiation from sibling. The description does not adequately compensate for the schema's lack of description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It specifies the 'topic' parameter implicitly ('based on a topic') but does not explain the 'max_results' parameter or any constraints/format for either parameter. The description adds little meaning beyond the schema's bare parameter names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Search for papers on arXiv based on a topic and store their information', which specifies the verb (search), resource (papers on arXiv), and action (store). However, it does not explicitly distinguish from the sibling tool 'extract_info', though the verb 'search' vs 'extract' provides implicit differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'extract_info', nor does it mention prerequisites, limitations, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v0.1.0- First observed
extract_info - First observed
search_papers
TDQS
Each tool has a clearly distinct purpose: one retrieves stored local information by paper ID, the other searches arXiv for papers by topic. No overlap or ambiguity.
Both tools follow a consistent verb_noun pattern: 'extract_info' and 'search_papers'. No mixed conventions.
With only 2 tools, the server feels under-scoped for a 'Research Server'. More tools like list, delete, or update would be expected, but the count is not extreme.
The surface covers search and retrieval by ID, but lacks essential operations like listing all stored papers, deleting, or updating entries, leaving notable gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Search arXiv, fetch paper metadata, and read full-text content.
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
arXiv MCP — preprint server search (free, no auth)
Related MCP Servers
- AlicenseBqualityNot gradedmaintenanceEnables AI assistants to search and retrieve academic papers from arXiv through MCP tools, supporting search by various criteria, detailed paper information, category browsing, and PDF content extraction.41292-
- AlicenseNot gradedqualityAmaintenanceSearch arXiv, fetch paper metadata, and read full-text content via MCP.6354Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables searching arXiv, fetching metadata, reading papers as section-aware Markdown, listing recent papers, and downloading PDFs via five MCP tools.232MIT
- AlicenseNot gradedqualityBmaintenanceEnables searching arXiv and downloading papers via MCP tools, no API key required.Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/akhileshvj/MCPChatbotForPapers'
If you have feedback or need assistance with the MCP directory API, please join our Discord server