Skip to main content
Glama
519,972 tools. Updated 2026-09-06 08:15

"Understanding Inference Models" matching MCP tools:

  • Retrieve all served vLLM models and LoRA adapters from a specified inference target, with adapters flagged, for governance-grade AIops monitoring and root-cause analysis.
    MIT
  • List available Claude Code CLI models, including identifiers, resolved models, descriptions, and supported effort or mode capabilities, without running an inference turn.
    MIT

Matching MCP Servers

Matching MCP Connectors

  • Search and use Hugging Face models, datasets, and spaces through inference, information retrieval, and authentication operations.
    Mozilla Public 2.0
  • Find available AI and voice models on Floe Inference, including their IDs, modalities, and context windows, to use with OpenAI-compatible endpoints.
    MIT
  • Generates detailed textual descriptions of one or more images for text-only models, enabling scene/UI understanding, OCR, comparison, and structured extraction from multi-image uploads.
    MIT
  • Identify running models and engine server details across vLLM, SGLang, and TGI. Provides engine-agnostic access to model IDs and server information from /v1/models or /info.
    MIT
  • Retrieve real-time health metrics for inference providers: health score, breaker state, uptime ratio, and recent failures.
    MIT
  • Load a model into memory for inference, with optional keep-alive duration and provider selection.
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International
  • List, load, or unload models on a local LLM server to switch active models in LM Studio without opening the GUI.
    Apache 2.0
  • Retrieves KV-cache utilization, prefix-cache hit rate, and preemption count to monitor inference cluster health and diagnose performance issues.
    MIT