Skip to main content
Glama
harinarayn

agent-trace-intelligence

by harinarayn

Agent Trace Intelligence MCP

Python 3.11+ MCP PyPI License MIT CI

Diagnose why your agent did what it did, and how to fix it.

The Problem

When an AI agent fails or behaves unexpectedly, existing tools tell you what happened: token counts, step logs, latency metrics. But they don't tell you why the agent made a wrong turn or how to fix it. Debugging agents means staring at raw traces and guessing.

Related MCP server: CI Investigator MCP

The Solution

Agent traces show you what happened. This tool tells you why it went wrong and what to change. Pass in any agent trace JSON and get root causes, scores, and a concrete fix back. Zero instrumentation required.

Use it alongside LangSmith, Arize Phoenix, and W&B Weave. When your observability stack surfaces a failure, this is where you go to diagnose it.

Works in Cursor, Claude Desktop, VS Code (Copilot MCP), and any stdio MCP client.


How It Fits

This tool explains why a single agent trace behaved the way it did.

It complements existing observability tools:

  • Azure Application Insights, AWS CloudWatch, GCP Cloud Trace: show what happened across runs

  • LangSmith, Arize Phoenix, W&B Weave: track agent behaviour over time

This tool answers a narrower question: why did this specific trace fail, and what exactly needs to change?


Install

# With uv (recommended)
uv add agent-trace-intelligence

# Or pip
pip install agent-trace-intelligence

Quick Start: Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "agent-trace-intelligence": {
      "command": "uv",
      "args": ["run", "agent-trace-intelligence"],
      "env": {
        "AZURE_AI_API_KEY": "your-azure-ai-foundry-key",
        "AZURE_AI_API_BASE": "https://your-resource.cognitiveservices.azure.com/",
        "JUDGE_MODEL": "azure_ai/claude-opus-4-6"
      }
    }
  }
}

No Azure? Use OpenAI instead:

{
  "env": {
    "JUDGE_MODEL": "gpt-4o-mini",
    "OPENAI_API_KEY": "sk-..."
  }
}

Tools Reference

Tool

Description

API Key?

Speed

judge_trace

Root cause analysis, 4-dimension scoring, grade, verdict, plain-English explanation

Required

~3-5s

trace_breakdown

Step-by-step scoring with flags (REDUNDANT_TOOL_CALL, REASONING_GAP, etc.)

Required

~3-5s

efficiency_score

Deterministic token/latency/redundancy analysis

Not required

Instant

judge_trace

Input:

{
  "trace": "<JSON string of AgentTrace>",
  "goal": "optional: override the goal stated in the trace"
}

Output:

{
  "overall_score": 0.82,
  "grade": "B",
  "verdict": "needs optimisation",
  "dimension_scores": {
    "goal_completion": 0.9,
    "reasoning_clarity": 0.8,
    "tool_usage": 0.75,
    "output_quality": 0.83
  },
  "summary": "Agent completed the goal but made one redundant search call",
  "root_causes": [
    "Unnecessary second search call at step 4 caused token inflation. Result was already available from step 2",
    "Agent did not validate tool output before proceeding to the next step"
  ],
  "strengths": ["Clear reasoning steps", "Correct initial tool selection"],
  "weaknesses": ["Redundant tool call on step 4", "Incomplete final output"],
  "recommendation": "Remove duplicate search call at step 4. Saves ~400 tokens",
  "explain_like_im_5": "The agent searched the internet twice for the same thing when it only needed to do it once, which wasted time and money.",
  "confidence": "high"
}

Verdict values: "production-ready" | "needs optimisation" | "broken"

trace_breakdown

Per-step scoring with flags:

  • REDUNDANT_TOOL_CALL: same tool called with same/similar input

  • HALLUCINATED_TOOL: tool referenced that doesn't exist in the trace

  • REASONING_GAP: response doesn't follow from previous tool output

  • GOAL_DRIFT: agent deviates from the original goal

  • PREMATURE_STOP: agent stopped before completing the goal

efficiency_score

No API key required. Deterministic analysis of:

  • Token usage (total, per-step, rating: good/moderate/high)

  • Tool redundancy (redundant calls, failed calls, redundancy rate)

  • Latency (total ms, slowest step/tool, rating: fast/acceptable/slow)

  • overall_efficiency_score (0.0-1.0 weighted composite)


AgentTrace Schema

All tools accept a trace JSON conforming to this schema:

{
  "trace_id": "optional",
  "agent_name": "optional",
  "goal": "What was the agent trying to do?",
  "model": "gpt-4o",
  "total_tokens": 820,
  "total_latency_ms": 3200,
  "final_output": "The agent's final response",
  "steps": [
    {
      "step_number": 1,
      "role": "user",
      "content": "User message"
    },
    {
      "step_number": 2,
      "role": "assistant",
      "content": "I'll search for that.",
      "token_count": 120
    },
    {
      "step_number": 3,
      "role": "tool",
      "tool_call": {
        "tool_name": "web_search",
        "input": {"query": "..."},
        "output": "Search results...",
        "latency_ms": 1200,
        "error": null
      }
    }
  ]
}

All fields except steps are optional. Works with whatever you can provide.


Format Adapters

Optional helpers to convert native framework traces to AgentTrace format:

from agent_trace_intelligence.formats import (
    adapt_langchain,      # LangChain callback handler output
    adapt_openai_agents,  # OpenAI Agents SDK RunStep objects
    adapt_autogen,        # AutoGen message history
    adapt_maf,            # MAF GA 1.0 (OpenTelemetry GenAI spans)
)

# Convert and pass directly to any tool
trace = adapt_langchain(raw_langchain_output)

These are convenience helpers. The tools accept any valid AgentTrace JSON regardless of framework.

Adapter

Framework

Status

adapt_langchain

LangChain callback handler / LangSmith export

v1

adapt_openai_agents

OpenAI Agents SDK (RunStep objects)

v1

adapt_autogen

AutoGen legacy (pyautogen) message history

v1

adapt_maf

Microsoft Agent Framework GA 1.0 (OTel spans)

v1


Model Support & Cost Guidance

Configure via JUDGE_MODEL env var. Zero code change required.

Use Case

Recommended Model

Cost

Best quality (default)

azure_ai/claude-opus-4-6

~$0.015/trace

Fast Azure alternative

azure_ai/gpt-4.1

~$0.008/trace

Open source / no Azure

gpt-4o-mini

~$0.002/trace

CI/CD batch evaluation

gpt-4o-mini

< $0.01/trace

Anthropic direct

claude-haiku-4-5-20251001

~$0.001/trace

For CI/CD use: Set JUDGE_MODEL=gpt-4o-mini to keep costs under $0.01 per trace. For interactive debugging, azure_ai/claude-opus-4-6 gives the best root cause reasoning.

efficiency_score is always free. No model call, no API key.


Future Direction (v2)

  • Pattern detection across multiple traces to surface recurring failure modes

  • Batch trace analysis for CI/CD quality gates

  • Enterprise governance signals to flag traces that violate defined agent policies

  • SSE transport for enterprise internal MCP deployment

  • Connectors to pull traces directly from observability platforms (Azure App Insights, AWS CloudWatch, GCP Cloud Trace, LangSmith). Contributions welcome.


License

MIT. See LICENSE

Available Tools

3 tools
efficiency_scoreA

Deterministic efficiency analysis of token usage, tool redundancy, and latency. No API key required — runs instantly.

ParametersJSON Schema
NameRequiredDescriptionDefault
traceYesAgent trace as a JSON string conforming to AgentTrace schema

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavioral traits. It explains that the analysis is deterministic, requires no API key, and runs instantly, which is useful. However, it does not disclose whether the operation is read-only, potential side effects, or what happens on invalid input, leaving some ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the purpose and immediately adds key behavioral cues (deterministic, no API key, instant). Every word contributes value with no redundancy or irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter and no output schema. The description explains the purpose and key benefits but does not clarify the return format or structure beyond calling it an 'efficiency analysis'. Given the absence of an output schema, this is a notable gap, though the name 'efficiency_score' partly compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage for the single parameter 'trace', describing it as a JSON string conforming to the AgentTrace schema. The description adds no parameter-specific meaning beyond the schema, so it relies entirely on the schema's documentation. The baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a specific purpose: deterministic efficiency analysis of token usage, tool redundancy, and latency. It distinguishes itself from siblings by focusing on efficiency metrics rather than judgment (judge_trace) or decomposition (trace_breakdown), making it easy for an agent to select when such analysis is needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for usage: it is deterministic, requires no API key, and runs instantly. These traits suggest appropriate scenarios (e.g., quick, low-cost analysis). However, it does not explicitly state when not to use it or mention alternative tools, so it falls short of the highest level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

judge_traceB

Diagnoses an agent trace — identifies root causes of failure, scores performance across four dimensions, and suggests a concrete fix. Returns verdict, grade, and plain-English explanation.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoOptional: override the goal stated in the trace
traceYesAgent trace as a JSON string conforming to AgentTrace schema

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses main outputs (verdict, grade, explanation) and that it scores across four dimensions, but does not mention error handling, side effects, or what the four dimensions are. This is adequate but not rich context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that wastes no words, efficiently covering purpose, method, and outputs. It is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool without an output schema, the description covers the essential aspects: what it does, what it returns, and that it evaluates four dimensions. It omits the specific dimensions and lacks usage guidance, but the core information for invoking the tool is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both parameters described ('Agent trace as a JSON string' and 'Optional: override the goal stated in the trace'). The description adds no parameter-specific detail beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Diagnoses an agent trace' and specifies outputs (verdict, grade, plain-English explanation) and that it 'scores performance across four dimensions.' However, it does not explicitly differentiate from sibling tools like trace_breakdown or efficiency_score, so it is clear but not fully distinguishing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It only explains what the tool does, not the contexts that favor it over trace_breakdown or efficiency_score, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_breakdownB

Step-by-step scoring of every agent decision. Flags issues like redundant tool calls, reasoning gaps, and goal drift.

ParametersJSON Schema
NameRequiredDescriptionDefault
traceYesAgent trace as a JSON string conforming to AgentTrace schema

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the transparency burden. It discloses that the tool produces step-by-step scores and flags specific issue types (redundant calls, reasoning gaps, goal drift), which gives some behavioral insight. However, it does not describe output format, side effects (likely none), or any limitations, leaving moderate ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core function and then giving concrete examples of what it flags. Every word adds value, with no redundancy or filler. It is appropriately sized for the tool's simple interface (one parameter).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one parameter and no output schema, so the description needs to convey the tool's behavior and return value. It explains what the tool does (scoring and flagging) but does not specify the output structure, scoring scale, or how the results are presented. It is adequate but leaves gaps for an agent trying to use the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides a complete description for the only parameter ('trace') with 100% coverage. The tool description adds no additional meaning about the parameter. Since schema coverage is high, the baseline of 3 applies, and the description does not enhance understanding of the parameter beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's function: 'Step-by-step scoring of every agent decision' with a specific verb ('scoring') and resource ('agent decision'). It goes further to list concrete issues it flags (redundant tool calls, reasoning gaps, goal drift), which distinguishes it from the sibling tools 'judge_trace' and 'efficiency_score' despite not naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given on when to use this tool versus alternatives. The description implies it is for detailed per-decision analysis, but it does not state exclusions or mention sibling tools. An agent would not know whether to choose this over 'judge_trace' or 'efficiency_score' based on the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 3 tool updatesv0.1.2
    • First observedefficiency_score
    • First observedjudge_trace
    • First observedtrace_breakdown

TDQS

A3.5/5.0
Disambiguation4/5

The tools are largely distinct: judge_trace provides an overall verdict, trace_breakdown offers step-by-step scoring, and efficiency_score focuses on efficiency metrics. However, judge_trace and trace_breakdown both provide performance scores, which could cause some confusion about which to use for a given task.

Naming Consistency2/5

The naming pattern is inconsistent: judge_trace follows a verb_noun convention while trace_breakdown and efficiency_score use noun_noun. Although all names use snake_case, the lack of a consistent pattern makes it harder to predict tool names.

Tool Count5/5

With only three tools, the server is well-scoped for its niche purpose of agent trace analysis. Each tool addresses a distinct aspect (overall judgment, detailed breakdown, and efficiency), so none feels redundant.

Completeness4/5

The server covers the core analysis lifecycle: holistic diagnosis, step-by-step scoring, and efficiency measurement. Minor gaps exist, such as the lack of tools for comparing multiple traces or retrieving raw trace data, but these are not critical to the primary function.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/harinarayn/agent-trace-intelligence'

If you have feedback or need assistance with the MCP directory API, please join our Discord server