agent-trace-intelligence
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-trace-intelligenceanalyze trace for root causes"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Trace Intelligence MCP
Diagnose why your agent did what it did, and how to fix it.
The Problem
When an AI agent fails or behaves unexpectedly, existing tools tell you what happened: token counts, step logs, latency metrics. But they don't tell you why the agent made a wrong turn or how to fix it. Debugging agents means staring at raw traces and guessing.
Related MCP server: CI Investigator MCP
The Solution
Agent traces show you what happened. This tool tells you why it went wrong and what to change. Pass in any agent trace JSON and get root causes, scores, and a concrete fix back. Zero instrumentation required.
Use it alongside LangSmith, Arize Phoenix, and W&B Weave. When your observability stack surfaces a failure, this is where you go to diagnose it.
Works in Cursor, Claude Desktop, VS Code (Copilot MCP), and any stdio MCP client.
How It Fits
This tool explains why a single agent trace behaved the way it did.
It complements existing observability tools:
Azure Application Insights, AWS CloudWatch, GCP Cloud Trace: show what happened across runs
LangSmith, Arize Phoenix, W&B Weave: track agent behaviour over time
This tool answers a narrower question: why did this specific trace fail, and what exactly needs to change?
Install
# With uv (recommended)
uv add agent-trace-intelligence
# Or pip
pip install agent-trace-intelligenceQuick Start: Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"agent-trace-intelligence": {
"command": "uv",
"args": ["run", "agent-trace-intelligence"],
"env": {
"AZURE_AI_API_KEY": "your-azure-ai-foundry-key",
"AZURE_AI_API_BASE": "https://your-resource.cognitiveservices.azure.com/",
"JUDGE_MODEL": "azure_ai/claude-opus-4-6"
}
}
}
}No Azure? Use OpenAI instead:
{
"env": {
"JUDGE_MODEL": "gpt-4o-mini",
"OPENAI_API_KEY": "sk-..."
}
}Tools Reference
Tool | Description | API Key? | Speed |
| Root cause analysis, 4-dimension scoring, grade, verdict, plain-English explanation | Required | ~3-5s |
| Step-by-step scoring with flags (REDUNDANT_TOOL_CALL, REASONING_GAP, etc.) | Required | ~3-5s |
| Deterministic token/latency/redundancy analysis | Not required | Instant |
judge_trace
Input:
{
"trace": "<JSON string of AgentTrace>",
"goal": "optional: override the goal stated in the trace"
}Output:
{
"overall_score": 0.82,
"grade": "B",
"verdict": "needs optimisation",
"dimension_scores": {
"goal_completion": 0.9,
"reasoning_clarity": 0.8,
"tool_usage": 0.75,
"output_quality": 0.83
},
"summary": "Agent completed the goal but made one redundant search call",
"root_causes": [
"Unnecessary second search call at step 4 caused token inflation. Result was already available from step 2",
"Agent did not validate tool output before proceeding to the next step"
],
"strengths": ["Clear reasoning steps", "Correct initial tool selection"],
"weaknesses": ["Redundant tool call on step 4", "Incomplete final output"],
"recommendation": "Remove duplicate search call at step 4. Saves ~400 tokens",
"explain_like_im_5": "The agent searched the internet twice for the same thing when it only needed to do it once, which wasted time and money.",
"confidence": "high"
}Verdict values: "production-ready" | "needs optimisation" | "broken"
trace_breakdown
Per-step scoring with flags:
REDUNDANT_TOOL_CALL: same tool called with same/similar inputHALLUCINATED_TOOL: tool referenced that doesn't exist in the traceREASONING_GAP: response doesn't follow from previous tool outputGOAL_DRIFT: agent deviates from the original goalPREMATURE_STOP: agent stopped before completing the goal
efficiency_score
No API key required. Deterministic analysis of:
Token usage (total, per-step, rating: good/moderate/high)
Tool redundancy (redundant calls, failed calls, redundancy rate)
Latency (total ms, slowest step/tool, rating: fast/acceptable/slow)
overall_efficiency_score(0.0-1.0 weighted composite)
AgentTrace Schema
All tools accept a trace JSON conforming to this schema:
{
"trace_id": "optional",
"agent_name": "optional",
"goal": "What was the agent trying to do?",
"model": "gpt-4o",
"total_tokens": 820,
"total_latency_ms": 3200,
"final_output": "The agent's final response",
"steps": [
{
"step_number": 1,
"role": "user",
"content": "User message"
},
{
"step_number": 2,
"role": "assistant",
"content": "I'll search for that.",
"token_count": 120
},
{
"step_number": 3,
"role": "tool",
"tool_call": {
"tool_name": "web_search",
"input": {"query": "..."},
"output": "Search results...",
"latency_ms": 1200,
"error": null
}
}
]
}All fields except steps are optional. Works with whatever you can provide.
Format Adapters
Optional helpers to convert native framework traces to AgentTrace format:
from agent_trace_intelligence.formats import (
adapt_langchain, # LangChain callback handler output
adapt_openai_agents, # OpenAI Agents SDK RunStep objects
adapt_autogen, # AutoGen message history
adapt_maf, # MAF GA 1.0 (OpenTelemetry GenAI spans)
)
# Convert and pass directly to any tool
trace = adapt_langchain(raw_langchain_output)These are convenience helpers. The tools accept any valid AgentTrace JSON regardless of framework.
Adapter | Framework | Status |
| LangChain callback handler / LangSmith export | v1 |
| OpenAI Agents SDK (RunStep objects) | v1 |
| AutoGen legacy ( | v1 |
| Microsoft Agent Framework GA 1.0 (OTel spans) | v1 |
Model Support & Cost Guidance
Configure via JUDGE_MODEL env var. Zero code change required.
Use Case | Recommended Model | Cost |
Best quality (default) |
| ~$0.015/trace |
Fast Azure alternative |
| ~$0.008/trace |
Open source / no Azure |
| ~$0.002/trace |
CI/CD batch evaluation |
| < $0.01/trace |
Anthropic direct |
| ~$0.001/trace |
For CI/CD use: Set JUDGE_MODEL=gpt-4o-mini to keep costs under $0.01 per trace. For interactive debugging, azure_ai/claude-opus-4-6 gives the best root cause reasoning.
efficiency_score is always free. No model call, no API key.
Future Direction (v2)
Pattern detection across multiple traces to surface recurring failure modes
Batch trace analysis for CI/CD quality gates
Enterprise governance signals to flag traces that violate defined agent policies
SSE transport for enterprise internal MCP deployment
Connectors to pull traces directly from observability platforms (Azure App Insights, AWS CloudWatch, GCP Cloud Trace, LangSmith). Contributions welcome.
License
MIT. See LICENSE
Available Tools
3 toolsefficiency_scoreA
Deterministic efficiency analysis of token usage, tool redundancy, and latency. No API key required — runs instantly.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes | Agent trace as a JSON string conforming to AgentTrace schema |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavioral traits. It explains that the analysis is deterministic, requires no API key, and runs instantly, which is useful. However, it does not disclose whether the operation is read-only, potential side effects, or what happens on invalid input, leaving some ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the purpose and immediately adds key behavioral cues (deterministic, no API key, instant). Every word contributes value with no redundancy or irrelevant detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema. The description explains the purpose and key benefits but does not clarify the return format or structure beyond calling it an 'efficiency analysis'. Given the absence of an output schema, this is a notable gap, though the name 'efficiency_score' partly compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage for the single parameter 'trace', describing it as a JSON string conforming to the AgentTrace schema. The description adds no parameter-specific meaning beyond the schema, so it relies entirely on the schema's documentation. The baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific purpose: deterministic efficiency analysis of token usage, tool redundancy, and latency. It distinguishes itself from siblings by focusing on efficiency metrics rather than judgment (judge_trace) or decomposition (trace_breakdown), making it easy for an agent to select when such analysis is needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for usage: it is deterministic, requires no API key, and runs instantly. These traits suggest appropriate scenarios (e.g., quick, low-cost analysis). However, it does not explicitly state when not to use it or mention alternative tools, so it falls short of the highest level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
judge_traceB
Diagnoses an agent trace — identifies root causes of failure, scores performance across four dimensions, and suggests a concrete fix. Returns verdict, grade, and plain-English explanation.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | Optional: override the goal stated in the trace | |
| trace | Yes | Agent trace as a JSON string conforming to AgentTrace schema |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses main outputs (verdict, grade, explanation) and that it scores across four dimensions, but does not mention error handling, side effects, or what the four dimensions are. This is adequate but not rich context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that wastes no words, efficiently covering purpose, method, and outputs. It is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool without an output schema, the description covers the essential aspects: what it does, what it returns, and that it evaluates four dimensions. It omits the specific dimensions and lacks usage guidance, but the core information for invoking the tool is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described ('Agent trace as a JSON string' and 'Optional: override the goal stated in the trace'). The description adds no parameter-specific detail beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Diagnoses an agent trace' and specifies outputs (verdict, grade, plain-English explanation) and that it 'scores performance across four dimensions.' However, it does not explicitly differentiate from sibling tools like trace_breakdown or efficiency_score, so it is clear but not fully distinguishing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It only explains what the tool does, not the contexts that favor it over trace_breakdown or efficiency_score, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_breakdownB
Step-by-step scoring of every agent decision. Flags issues like redundant tool calls, reasoning gaps, and goal drift.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes | Agent trace as a JSON string conforming to AgentTrace schema |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the transparency burden. It discloses that the tool produces step-by-step scores and flags specific issue types (redundant calls, reasoning gaps, goal drift), which gives some behavioral insight. However, it does not describe output format, side effects (likely none), or any limitations, leaving moderate ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core function and then giving concrete examples of what it flags. Every word adds value, with no redundancy or filler. It is appropriately sized for the tool's simple interface (one parameter).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one parameter and no output schema, so the description needs to convey the tool's behavior and return value. It explains what the tool does (scoring and flagging) but does not specify the output structure, scoring scale, or how the results are presented. It is adequate but leaves gaps for an agent trying to use the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides a complete description for the only parameter ('trace') with 100% coverage. The tool description adds no additional meaning about the parameter. Since schema coverage is high, the baseline of 3 applies, and the description does not enhance understanding of the parameter beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's function: 'Step-by-step scoring of every agent decision' with a specific verb ('scoring') and resource ('agent decision'). It goes further to list concrete issues it flags (redundant tool calls, reasoning gaps, goal drift), which distinguishes it from the sibling tools 'judge_trace' and 'efficiency_score' despite not naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given on when to use this tool versus alternatives. The description implies it is for detailed per-decision analysis, but it does not state exclusions or mention sibling tools. An agent would not know whether to choose this over 'judge_trace' or 'efficiency_score' based on the description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v0.1.2- First observed
efficiency_score - First observed
judge_trace - First observed
trace_breakdown
TDQS
The tools are largely distinct: judge_trace provides an overall verdict, trace_breakdown offers step-by-step scoring, and efficiency_score focuses on efficiency metrics. However, judge_trace and trace_breakdown both provide performance scores, which could cause some confusion about which to use for a given task.
The naming pattern is inconsistent: judge_trace follows a verb_noun convention while trace_breakdown and efficiency_score use noun_noun. Although all names use snake_case, the lack of a consistent pattern makes it harder to predict tool names.
With only three tools, the server is well-scoped for its niche purpose of agent trace analysis. Each tool addresses a distinct aspect (overall judgment, detailed breakdown, and efficiency), so none feels redundant.
The server covers the core analysis lifecycle: holistic diagnosis, step-by-step scoring, and efficiency measurement. Minor gaps exist, such as the lack of tools for comparing multiple traces or retrieving raw trace data, but these are not critical to the primary function.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
AI agent observability for production traces, natural-language insights, and improvement loops.
Find your AI agent's likely failure mode, get runtime settings, and clarify ambiguous prompts.
Lints + auto-fixes how AI coding agents discover any new product. 24 rules, 6 tools, score 0-100.
Data + AI observability — monitor and troubleshoot production-grade agents and the context they use.
Related MCP Servers
- AlicenseBqualityBmaintenanceSupercharges AI-assisted debugging of Playwright tests by parsing trace files to extract failures, action history, network logs, screenshots, and suggesting fixes.614MIT
- AlicenseAqualityAmaintenanceProvides tools to analyze and debug GitHub Actions CI failures, including summarizing failures, detecting flaky tests, and suggesting fixes.10171ISC
- AlicenseNot gradedqualityCmaintenanceEnables recording and analyzing AI agent execution traces, including event logging, metric computation, loop detection, and JSON export for debugging agent behavior.MIT

Pisama MCP Serverofficial
AlicenseNot gradedqualityBmaintenanceEnables analysis of AI agent traces to detect and fix failures using heuristic detectors, with no LLM calls required.1MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/harinarayn/agent-trace-intelligence'
If you have feedback or need assistance with the MCP directory API, please join our Discord server