SlimContext MCP Server
Uses OpenAI's API to generate AI-powered summaries of chat history, compressing conversations while preserving context using models like gpt-4o-mini.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SlimContext MCP Serversummarize this long chat history to fit within 4000 tokens"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SlimContext MCP Server
A Model Context Protocol (MCP) server that wraps the SlimContext library, providing AI chat history compression tools for MCP-compatible clients.
Overview
SlimContext MCP Server exposes two powerful compression strategies as MCP tools:
trim_messages- Token-based compression that removes oldest messages when exceeding token thresholdssummarize_messages- AI-powered compression using OpenAI to create concise summaries
Related MCP server: PromptThrift MCP
Installation
npm install -g slimcontext-mcp-server
# or
pnpm add -g slimcontext-mcp-serverDevelopment
# Clone and setup
git clone <repository>
cd slimcontext-mcp-server
pnpm install
# Build
pnpm build
# Run in development
pnpm dev
# Type checking
pnpm typecheckConfiguration
MCP Client Setup
Add to your MCP client configuration:
{
"mcpServers": {
"slimcontext": {
"command": "npx",
"args": ["-y", "slimcontext-mcp-server"]
}
}
}Environment Variables
OPENAI_API_KEY: OpenAI API key for summarization (optional, can be passed as tool parameter)
Tools
trim_messages
Compresses chat history using token-based trimming strategy.
Parameters:
messages(required): Array of chat messagesmaxModelTokens(optional): Maximum model token context window (default: 8192)thresholdPercent(optional): Percentage threshold to trigger compression 0-1 (default: 0.7)minRecentMessages(optional): Minimum recent messages to preserve (default: 2)
Example:
{
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Hello!" },
{ "role": "assistant", "content": "Hi there! How can I help you today?" },
{ "role": "user", "content": "Tell me about AI." }
],
"maxModelTokens": 4000,
"thresholdPercent": 0.8,
"minRecentMessages": 2
}Response:
{
"success": true,
"original_message_count": 4,
"compressed_message_count": 3,
"messages_removed": 1,
"compression_ratio": 0.75,
"compressed_messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "assistant", "content": "Hi there! How can I help you today?" },
{ "role": "user", "content": "Tell me about AI." }
]
}summarize_messages
Compresses chat history using AI-powered summarization strategy.
Parameters:
messages(required): Array of chat messagesmaxModelTokens(optional): Maximum model token context window (default: 8192)thresholdPercent(optional): Percentage threshold to trigger compression 0-1 (default: 0.7)minRecentMessages(optional): Minimum recent messages to preserve (default: 4)openaiApiKey(optional): OpenAI API key (can also use OPENAI_API_KEY env var)openaiModel(optional): OpenAI model for summarization (default: 'gpt-4o-mini')customPrompt(optional): Custom summarization prompt
Example:
{
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "I want to build a web scraper." },
{
"role": "assistant",
"content": "I can help you build a web scraper! What programming language would you prefer?"
},
{ "role": "user", "content": "Python please." },
{
"role": "assistant",
"content": "Great choice! For Python web scraping, I recommend using requests and BeautifulSoup..."
},
{ "role": "user", "content": "Can you show me a simple example?" }
],
"maxModelTokens": 4000,
"thresholdPercent": 0.6,
"minRecentMessages": 2,
"openaiModel": "gpt-4o-mini"
}Response:
{
"success": true,
"original_message_count": 6,
"compressed_message_count": 4,
"messages_removed": 2,
"summary_generated": true,
"compression_ratio": 0.67,
"compressed_messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{
"role": "system",
"content": "The user expressed interest in building a web scraper and requested help with Python. The assistant recommended using requests and BeautifulSoup libraries for Python web scraping."
},
{
"role": "assistant",
"content": "Great choice! For Python web scraping, I recommend using requests and BeautifulSoup..."
},
{ "role": "user", "content": "Can you show me a simple example?" }
]
}Message Format
Both tools expect messages in SlimContext format:
interface SlimContextMessage {
role: 'system' | 'user' | 'assistant' | 'tool' | 'human';
content: string;
}Error Handling
All tools return structured error responses:
{
"success": false,
"error": "Error message description",
"error_type": "SlimContextError" | "OpenAIError" | "UnknownError"
}Common error scenarios:
Missing OpenAI API key for summarization
Invalid message format
OpenAI API rate limits or errors
Invalid parameter values
Token Estimation
SlimContext uses a simple heuristic for token estimation: Math.ceil(content.length / 4) + 2. This provides a reasonable approximation for most use cases. For more accurate token counting, you would need to implement a custom token estimator in your client application.
Compression Strategies
Trimming Strategy
Preserves all system messages
Preserves the most recent N messages
Removes oldest non-system messages until under token threshold
Fast and deterministic
No external API dependencies
Summarization Strategy
Preserves all system messages
Preserves the most recent N messages
Summarizes middle portion of conversation using AI
Creates contextually rich summaries
Requires OpenAI API access
License
MIT
Contributing
Fork the repository
Create a feature branch
Make your changes
Add tests for new functionality
Submit a pull request
Related
SlimContext - The underlying compression library
Model Context Protocol - The protocol specification
MCP SDK - TypeScript SDK for MCP
Available Tools
2 toolssummarize_messagesA
Compress chat message history using AI-powered summarization strategy. Creates concise summaries of older messages while preserving system messages and recent context.
| Name | Required | Description | Default |
|---|---|---|---|
| messages | Yes | Array of chat messages to compress | |
| maxModelTokens | No | Model's maximum token context window | |
| thresholdPercent | No | Percentage threshold to trigger compression (0-1) | |
| minRecentMessages | No | Minimum recent messages to always preserve | |
| openaiApiKey | No | OpenAI API key (can also be set via OPENAI_API_KEY environment variable) | |
| openaiModel | No | OpenAI model to use for summarization | gpt-4o-mini |
| customPrompt | No | Custom prompt for summarization (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the tool's strategy ('AI-powered summarization'), what gets preserved ('system messages and recent context'), and what gets compressed ('older messages'), but doesn't mention rate limits, authentication requirements (though API key parameter hints at this), or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences that each earn their place: the first states the core functionality, the second explains the preservation strategy. No wasted words, well-structured, and front-loaded with the essential purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 7 parameters and no output schema, the description provides good context about the summarization approach and preservation logic. However, it doesn't explain what the output looks like (summary format) or potential limitations, leaving some gaps in completeness despite the strong schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all 7 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema, maintaining the baseline score of 3 for adequate but not enhanced parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('compress', 'creates concise summaries') and resources ('chat message history'), and distinguishes it from the sibling tool 'trim_messages' by specifying it uses 'AI-powered summarization strategy' rather than simple trimming.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('compress chat message history', 'creates concise summaries of older messages'), but doesn't explicitly state when NOT to use it or mention the sibling tool 'trim_messages' as an alternative for different compression needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trim_messagesA
Compress chat message history using token-based trimming strategy. Removes oldest non-system messages when token count exceeds threshold while preserving system messages and recent context.
| Name | Required | Description | Default |
|---|---|---|---|
| messages | Yes | Array of chat messages to compress | |
| maxModelTokens | No | Model's maximum token context window | |
| thresholdPercent | No | Percentage threshold to trigger compression (0-1) | |
| minRecentMessages | No | Minimum recent messages to always preserve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses key behavioral traits: the compression strategy (removes oldest non-system messages), preservation rules (preserves system messages and recent context), and triggering condition (token count exceeds threshold). However, it doesn't mention performance characteristics, error conditions, or what happens when compression isn't possible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with zero waste. The first sentence states the core purpose, the second explains the specific algorithm. Every word earns its place, and the most important information (what the tool does) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, 100% schema coverage, but no annotations and no output schema, the description provides adequate but not complete context. It explains the compression algorithm well but doesn't describe the return value format or error handling. The description compensates somewhat for the lack of annotations by explaining behavioral aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond what's in the schema - it mentions 'token-based trimming' which relates to the parameters but doesn't explain their interactions or provide additional semantic context. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'compress' and resource 'chat message history' with specific strategy details ('token-based trimming', 'removes oldest non-system messages'). It distinguishes from sibling 'summarize_messages' by focusing on compression rather than summarization.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('when token count exceeds threshold'), but doesn't explicitly state when NOT to use it or mention the sibling tool 'summarize_messages' as an alternative. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
- First observed
summarize_messages - First observed
trim_messages
TDQS
The two tools have clearly distinct purposes: summarize_messages uses AI-powered summarization to compress history by creating concise summaries, while trim_messages uses token-based trimming to remove oldest messages when exceeding thresholds. There is no overlap in their approaches, making them easily distinguishable.
Both tools follow a consistent verb_noun pattern with underscore separation: summarize_messages and trim_messages. The naming is predictable and readable, with no deviations in style or convention.
With only 2 tools, the server feels thin for a context management domain, as it lacks operations like retrieval, update, or deletion of summaries/trims. However, the tools cover compression strategies adequately for a minimal scope.
The server is severely incomplete for context management; it only offers compression methods (summarization and trimming) but lacks any tools to retrieve, modify, or manage the compressed contexts, leaving agents with no way to access or update the results of these operations.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Memory for deep conversational context across any platform
- AmberOAuthcom.ambermem
Long-term memory for AI assistants. Hybrid retrieval, query expansion, auto-topics.
Provide your AI coding tools with token-efficient access to up-to-date technical documentation for…
Your portable context layer — load it into any AI.
Related MCP Servers
- AlicenseBqualityDmaintenanceProvides intelligent context management for AI development sessions, allowing users to track token usage, manage conversation context, and seamlessly restore context when reaching token limits.8172Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables 70-90% LLM API cost reduction by compressing conversation history via local Gemma 4 models or heuristics, featuring token counting, model routing, and pinned facts for preserving critical context.1MIT

compresh-mcpofficial
AlicenseNot gradedqualityCmaintenanceProvides production-grade context compression for LLM agent conversations with Q-protective ranking, epistemic markers, and semantic store, reducing token usage while preserving equivalence.3Business Source 1.1- AlicenseNot gradedqualityCmaintenanceProvides context compression via the tokenslim engine, enabling MCP hosts to reduce token usage while preserving key information. Offers compress, retrieve, and stats tools for managing compressed content.Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/agentailor/slimcontext-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server