huoshui-pdf-translator
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@huoshui-pdf-translatortranslate Desktop/paper.pdf to Chinese"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Huoshui PDF Translator
Version: 0.1.0
Powered by: FastMCP & PDFMathTranslate-next
PyPI Package: huoshui-pdf-translator
An intelligent PDF translation assistant that specializes in academic papers with mathematical formulas. Built using the FastMCP framework and powered by PDFMathTranslate-next, it provides comprehensive translation capabilities with context-aware assistance.
🌟 Features
Core Translation Capabilities
📚 Academic Papers: Excellent handling of mathematical formulas and equations
🔬 Technical Documents: Preserves formatting and technical terminology
🌐 Multi-language Support: Auto-detection with Chinese ↔ English specialization
🎨 Layout Preservation: Maintains original PDF structure and formatting
Smart Assistant Features
🧠 Context-Aware Prompts: Multiple specialized prompts for different scenarios
🛠️ Tool Status Checking: Verify translation tool installation and availability
📊 PDF Analysis: Get detailed information about PDF files before translation
🔍 Flexible Path Handling: Support for both absolute and relative file paths
⚡ Progress Reporting: Real-time progress updates during translation
🚨 Intelligent Error Handling: Comprehensive error diagnosis and troubleshooting
MCP Features
📋 Resources: Translation capability listings and PDF file information
🎯 Tools: Translation, PDF analysis, and tool status checking
💬 Prompts: Role definitions, path guidance, options explanation, and error troubleshooting
🔒 Security: Safe path validation with system directory protection
Related MCP server: PDF2ZH MCP Server
🚀 Quick Start
Installation
From MCP Registry (Recommended)
This server is available in the Model Context Protocol Registry. Install it using your MCP client.
mcp-name: io.github.huoshuiai42/huoshui-pdf-translator
Using uvx
uvx huoshui-pdf-translatorClaude Desktop Setup
Add this to your Claude Desktop MCP configuration:
{
"mcpServers": {
"huoshui-pdf-translator": {
"command": "uvx",
"args": ["huoshui-pdf-translator"]
}
}
}Alternative Installation Methods
Via pipx:
pipx install huoshui-pdf-translatorVia UV tools:
uv tool install huoshui-pdf-translatorClaude Desktop config for UV tools:
{
"mcpServers": {
"huoshui-pdf-translator": {
"command": "uv",
"args": ["tool", "run", "huoshui-pdf-translator"]
}
}
}📖 Usage
First-Time Setup
Warm up (downloads fonts/models): Use
warm_up_translatortoolCheck status: Use
check_translation_tooltoolTranslate: Use
translate_pdftool with your PDF path
MCP Tools
translate_pdf
Translates PDF documents while preserving mathematical formulas and layout.
# Basic usage
translate_pdf(pdf_path="Desktop/paper.pdf")
# With custom output path
translate_pdf(
pdf_path="Documents/research.pdf",
output_path="Documents/translated/research_cn.pdf"
)pdf_get
Retrieves detailed information about a PDF file.
pdf_info = pdf_get(path="Desktop/document.pdf")
# Returns: PDFResource with path, size_bytes, page_countwarm_up_translator
Downloads required assets and models. Run this first to avoid timeouts.
warm_up_translator()
# Downloads fonts and models (~50MB) for faster subsequent translationscheck_translation_tool
Verifies PDFMathTranslate-next installation and status.
status = check_translation_tool()
# Returns: status, version, messageMCP Prompts
role_and_rules: Core identity and operational rulesexplain_pdf_paths: Help with file path specificationsexplain_translation_options: Available options and best practicestroubleshoot_translation_error: Error diagnosis and solutionsexplain_translation_result: Result explanation and next steps
File Path Examples
The assistant supports flexible path specifications:
# Absolute paths
/Users/john/Desktop/research.pdf
C:\Users\John\Documents\paper.pdf
# Relative to home directory
Desktop/research.pdf
Documents/papers/study.pdf
# Simple filenames (assumes home directory)
paper.pdf🎯 Translation Workflow
Install:
uvx huoshui-pdf-translatorSetup Claude Desktop: Add MCP configuration
Warm up: Run
warm_up_translatortool (first time only)Translate: Use
translate_pdfwith your PDF pathReview: Two files created (dual-language and Chinese-only)
⚡ Performance
First translation: 2-5 minutes (downloads fonts/models)
Subsequent translations: 30-60 seconds
File size limit: 200MB maximum
Cache size: ~50MB for fonts and models
🔍 Troubleshooting
Common Issues
Translation Tool Not Available
The tool automatically installs pdf2zh-next when needed. If issues occur:
# Check status
# Use check_translation_tool in Claude Desktop
# Manual install if needed
pip install pdf2zh-nextFirst Translation Timeout
# Run warmup first
# Use warm_up_translator tool in Claude DesktopPDF File Not Found
Verify file path is correct
Use absolute paths for clarity
Check file hasn't been moved or deleted
Network Issues
Ensure internet connection (required for first-time font downloads)
Check firewall settings
Error Diagnosis
The assistant provides intelligent error diagnosis with specific solutions for:
File not found errors
Invalid PDF files
Translation tool issues
Network connectivity problems
File size limitations
🛠️ Development
For Developers
Install from source:
git clone https://github.com/huoshuiai/huoshui-pdf-translator.git
cd huoshui-pdf-translator
uv sync
uv run python -m huoshui_pdf_translator.mainBuild and publish:
uv build
uv run twine upload dist/*Project Structure
huoshui-pdf-translator/
├── huoshui_pdf_translator/
│ ├── __init__.py # Package metadata
│ └── main.py # FastMCP server implementation
├── pyproject.toml # Package configuration
├── README.md # This file
└── LICENSE # Apache-2.0 license🔄 Updates
Update to latest version:
uvx install --upgrade huoshui-pdf-translator
# or
uv tool upgrade huoshui-pdf-translator🤝 Contributing
Contributions are welcome! Please:
Fork the repository
Create a feature branch
Make your changes
Add tests if applicable
Submit a pull request
📄 License
This project is licensed under the Apache-2.0 License. See the LICENSE file for details.
🙏 Acknowledgments
PDFMathTranslate-next: Core translation engine
FastMCP: Framework for intelligent assistant capabilities
Anthropic: MCP protocol and ecosystem
UV & PyPI: Modern Python packaging and distribution
Available Tools
4 toolscheck_translation_toolA
Checks if the PDFMathTranslate-next tool is properly installed and available.
Returns: Dictionary with status information about the translation tool
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a read-only check ('Checks') and specifies the return type ('Dictionary with status information'), but does not elaborate on error behavior, edge cases, or what the status keys mean. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first states the core action, and the second specifies the return type. Every word earns its place; there is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple check tool with no parameters and an output schema, the description adequately covers purpose and return type. It lacks usage guidance and detailed behavior, but these are not critical for this tool's simplicity, and the output schema presumably fills in the return structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema is trivially covered. Per the rubric, a 0-parameter tool gets a baseline of 4, and the description adds no unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether the PDFMathTranslate-next tool is properly installed and available, using a specific verb ('Checks') and identifying the exact resource. This distinguishes it from sibling tools like translate_pdf or warm_up_translator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention whether it should be called before translation, any prerequisites, or situations where it would be unnecessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_getA
Retrieves detailed information about a PDF file.
Args: path: PDF file path (absolute or relative to home directory)
Returns: PDF resource with detailed information
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | PDF file path (absolute or relative to home directory) |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | PDF file path (absolute or relative to home directory) |
| page_count | No | Number of pages in the PDF |
| size_bytes | Yes | File size in bytes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of conveying behavior. It correctly implies a read-only operation and mentions the return type, but 'detailed information' is vague, and no context is given about potential errors, permissions, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, front-loaded with the purpose, and contains no unnecessary filler. The Args/Returns structure is clear and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no enums, output schema present), the description adequately covers the operation and expected return. The output schema fills in any remaining details about the return structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the description simply repeats the schema's parameter description. No additional semantic detail is provided beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Retrieves') and resource ('detailed information about a PDF file'). It is easily distinguished from sibling tools like translate_pdf or check_translation_tool, which focus on translation rather than information retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The sibling tools suggest a translation workflow, but the description does not position pdf_get within that workflow or mention any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translate_pdfA
Translates a PDF document using PDFMathTranslate-next. Preserves mathematical formulas and layout.
Args: pdf_path: PDF file path (absolute or relative to home directory) output_path: Optional custom output path for the translated PDF
Returns: Dictionary with paths to translated files (dual and mono versions)
| Name | Required | Description | Default |
|---|---|---|---|
| pdf_path | Yes | PDF file path (absolute or relative to home directory) | |
| output_path | No | Optional output path for translated PDF |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses preservation of mathematical formulas and layout, plus the return structure (dual and mono versions). However, it does not mention prerequisites like needing a warmed-up translator, potential file overwrites, or other side effects, leaving significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, starting with the main purpose, followed by Args and Returns sections. It is front-loaded with the core action. Some redundancy exists because the Args section repeats schema descriptions, but overall it is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with two parameters and an output schema exists (as indicated by context signals), so the description does not need to explain return values in detail. It covers the core behavior and output. However, it omits mention of the warm_up_translator sibling as a prerequisite, which is a minor completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% because both pdf_path and output_path have descriptions that match the Args section. The description adds no new meaning beyond the schema; it merely restates the parameter definitions. Baseline 3 is appropriate since the schema fully documents the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool translates a PDF document via PDFMathTranslate-next, with preservation of formulas and layout. This specific verb+resource combination distinguishes it from sibling tools like pdf_get, warm_up_translator, and check_translation_tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used for translating PDFs, but does not explicitly state when to use it versus the sibling tools (e.g., whether warm_up_translator should be called first). There is no direct alternative or exclusion guidance, only an implied purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warm_up_translatorA
Warm up the PDF translator by downloading required assets and models. Run this first to avoid timeouts during actual translation.
Returns: Dictionary with warmup status information
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description appropriately discloses that it downloads assets/models and returns a status dictionary. It also implies a side effect of fetching remote resources, which is useful behavioral information for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the main action, followed by usage timing and return value. No wasted words, all sentences add value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential aspects: what it does, when to run it, and what it returns, and an output schema covers detailed return values. It could mention network requirements or re-run behavior, but for a simple warm-up tool, it's reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description doesn't need to elaborate on parameter meanings. The baseline of 4 applies, and adding no parameter information is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to warm up the PDF translator by downloading required assets and models. It distinguishes itself from sibling tools like translate_pdf and pdf_get by explicitly positioning it as a preparatory step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear guidance to run this first to avoid timeouts during actual translation, establishing the tool as a prerequisite. However, it does not explicitly mention when not to use it or alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v0.1.1- First observed
check_translation_tool - First observed
pdf_get - First observed
translate_pdf - First observed
warm_up_translator
TDQS
Each tool has a distinct purpose: pdf_get retrieves PDF info, warm_up_translator prepares the model, translate_pdf performs the translation, and check_translation_tool verifies installation. There is no overlap or ambiguity between them.
Three tools follow a clear verb_noun pattern (warm_up_translator, translate_pdf, check_translation_tool), but pdf_get reverses this to noun_verb, creating a minor inconsistency. The naming is otherwise clear and readable.
With 4 tools, the server is well-scoped for its purpose of PDF translation. Each tool covers a necessary step in the workflow, and the count is appropriate.
The toolset covers the complete workflow: checking installation, warming up models, translating PDFs, and retrieving PDF details. There are no obvious gaps in the provided functionality.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
AI/ML research papers from arXiv, DBLP, and HuggingFace
Search 8.5M scientific papers with LLM TLDRs, citations, linked entities, figures, and full text.
Search 340M+ academic papers — citation graphs, semantic similarity, and AI literature reviews.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA paper retrieval and content parsing tool based on arXiv, supporting paper search, PDF link retrieval, and content parsing functions, suitable for obtaining the latest papers in academic research and AI fields.2Apache 2.0
- AlicenseAqualityDmaintenanceAn MCP server that translates scientific PDF documents while preserving original formulas, charts, and layout. It utilizes an OpenAI-compatible backend to provide tools for document translation and language listing.23AGPL 3.0

bluente-translateofficial
AlicenseAqualityDmaintenanceTranslate your documents with formatting intact in 2 minutes61911MIT- AlicenseNot gradedqualityCmaintenanceConverts LaTeX math expressions to SVG. Enables AI assistants to render mathematical formulas as scalable vector graphics.59MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/huoshuiai42/huoshui-pdf-translator'
If you have feedback or need assistance with the MCP directory API, please join our Discord server