Popular Model Observability AI tools

11 category leaders in Model Observability, selected from the full directory.

Evidently

  • Open-source evals and monitoring for LLMs and ML models.
  • 100+ metrics, LLM judges, drift, and data-quality checks.
  • Reports and dashboards for tests and production.
  • Apache 2.0 library with a managed Evidently Cloud.

Langfuse

  • Open-source LLM observability: traces, costs, and evals.
  • Prompt versioning, datasets, and LLM-as-judge scoring.
  • OpenTelemetry, LangChain, and LlamaIndex integrations.
  • Part of ClickHouse; self-host or cloud.

LangSmith

  • LangChain-native tracing, datasets, and regression evals.
  • Debug chains and agents with full run trees.
  • Prompt iteration and online/offline evaluation workflows.
  • Built in for LangChain and LangGraph deployments.

Braintrust

  • Eval and logging platform for AI products and agents.
  • Scorer SDKs, prompt experiments, and CI-friendly tests.
  • Traces production runs and surfaces new behaviors.
  • Notion runs evals across 70 engineers on it.

Arize

  • Agent and LLM observability, evals, and improvement.
  • Tracing, drift, and root-cause tools for production.
  • Phoenix open-source tracing plus enterprise AX platform.
  • Covers classic ML models and GenAI agents in one stack.

Galileo

  • LLM and agent evaluation with production observability.
  • Luna eval models for low-latency hallucination checks.
  • RAG, safety, and security evals turned into guardrails.
  • Used by Cisco, MongoDB, and Writer.

Phoenix

  • Arize OSS for LLM tracing and evaluation.
  • OpenTelemetry-native; self-host or Arize cloud.
  • Embeddings analysis and RAG eval primitives.
  • Apache 2.0 with a large open-source install base.

Weave

  • W&B Weave for LLM traces and evaluations.
  • Ties GenAI debugging to Weights & Biases workflows.
  • Agent and production observability for ML teams.
  • Included with Weights & Biases plans.

Opik

  • Comet open-source platform for LLM and agent observability.
  • Tracing, evals, and debugging for production GenAI.
  • Self-host or run on Comet Cloud beside experiment tracking.
  • Also tracks Claude Code usage and spend for teams.

Ragas

  • Open-source metrics for RAG and LLM app quality.
  • Faithfulness, relevance, and context precision/recall.
  • Synthetic eval data and online production monitoring.
  • Integrated with LangChain, LlamaIndex, and LangSmith.

Raindrop

  • Monitors production agent trajectories for silent failures.
  • Catches hallucinations, tool misuse, and model-upgrade regressions.
  • Simulations replay real traffic against proposed changes.
  • Used by Vercel, Framer, and Clay; $50M raised.