Model Observability AI tools

Model Observability platforms provide comprehensive monitoring and analysis of AI model performance and behavior. These tools enable organizations to track model health, detect issues, and ensure reliable AI operations through continuous monitoring and analysis.

37 verified AI-first sites in Model Observability.

Last reviewed .

Latest changes

Arize

ML and LLM observability for production models and agents. Traces, drift, embeddings, and root-cause tools, plus Phoenix open-source tracing for RAG and GenAI apps.

Pricing: Free starter. Enterprise custom; see arize.com.

Fiddler

Control plane for enterprise agents. Agentic observability, continuous monitoring, enforceable policy guardrails, and auditable governance. Used by Mastercard, Nielsen, US Navy, Ally, and AIG.

Pricing: Enterprise pricing. Free guardrails tier available.

Evidently

Open-source AI evaluation and observability for LLM apps, agents, and predictive models. 100+ built-in metrics, LLM-as-judge evals, drift and data-quality checks, and dashboards for tests and production monitoring. Evidently Cloud adds collaboration and tracing.

Pricing: Open Source: Free. Enterprise: Custom pricing.

Galileo

Observability and evaluation platform for LLMs and AI agents. Features Luna models for low-latency evals, RAG/agent/safety/security evaluations, hallucination detection, and production guardrails. Turn offline evals into real-time monitoring. Trusted by Writer, Cisco, Clearwater Analytics, MongoDB.

Pricing: Free tier. Enterprise: Custom pricing.

Superwise

Agentic management platform. Sentinel is an AI gateway that observes traffic and enforces runtime guardrails; Superwise Chat is a governed workspace for business teams.

Pricing: Free Solo edition. Enterprise; contact Superwise.

Arthur

Enterprise agent discovery and governance. Finds agents across endpoints and clouds, traces reasoning and tool calls, and enforces runtime policies with audit logs.

Pricing: Enterprise; contact Arthur.

Basalt

AI assistant observability that analyzes Langfuse, Braintrust, or Datadog traces overnight and opens one validated PR per day. Detects failure patterns in real conversations and proposes prompt, tool, or system-instruction fixes tested against a golden dataset before merge.

Pricing: First 3 merged PRs free; pay per merged PR after.

Deepchecks

Validation and monitoring platform for ML models and LLMs. Features automated testing, data quality checks, and model behavior analysis. Includes drift detection, performance monitoring, and custom validation suites. Offers both open-source and enterprise solutions.

Pricing: Free open-source tier. Enterprise pricing available.

Langfuse

Open-source LLM and agent observability and evals, now part of ClickHouse. Traces prompts, tool calls, costs, and latency; manages prompt versions, datasets, and LLM-as-judge evaluations; integrates with OpenTelemetry, LangChain, LlamaIndex, and major model providers. Self-host or cloud.

Pricing: Self-host free. Cloud Hobby free. Core $29/month. Pro $199/month. Enterprise from $2,499/month.

Helicone

Open-source AI gateway and LLM observability. Proxy-based logging of requests, costs, and latency across providers with routing, caching, and fallbacks; processed 14T+ tokens for 16,000 organizations. Acquired by Mintlify in March 2026 and now in maintenance mode-live and patched, with no new features.

Pricing: Hobby free (10k requests). Pro $79/month. Team $799/month. Enterprise custom.

Openlayer

AI governance, testing, and observability platform that inventories every AI system in a company. Pre-deploy tests, production monitoring for drift, latency, and prompt failures, cost controls, and audit-ready governance for both ML models and LLM agents.

Pricing: Basic trial with 20k inferences/month. Enterprise custom.

DataFramer

AI workflow intelligence for LLM apps-trace failures, expert review loops, and calibrated judges tied to accuracy, adoption, and business outcomes. Databricks-partnered eval/ops layer rather than classic UI regression QA.

Pricing: Contact for pricing.

LangSmith

LangChain-native observability for LLM applications: distributed tracing over chains and tools, datasets for regression testing, online and offline evaluation, prompt iteration, and collaboration for ML/AI teams. Tight integration with LangChain and LangGraph; positions debugging and quality gates as first-class for production GenAI-not ad hoc logging.

Pricing: Developer tier; usage-based and enterprise; see site.

Braintrust

Evaluation and observability platform for AI products: logged traces, scorer SDKs, prompt and experiment comparison, and CI-friendly eval workflows. Emphasizes fast iteration on LLM quality with developer-centric UX and data ownership. Complements generic APM with task-specific LLM metrics and human/LLM-as-judge patterns.

Pricing: Free tier; team and enterprise; see site.

Patronus

AI research lab and evaluation company. Percival debugs agent traces and flags failure patterns; its core platform runs LLM tests, safety checks, and judges; newer work builds RL environments and Digital World Models for training frontier agents.

Pricing: Contact; enterprise; see site.

Portkey

AI gateway and control plane with routing to 1,600+ models, fallbacks, observability, guardrails, governance, and prompt management. Acquired by Palo Alto Networks in May 2026 and now sold as the Prisma AIRS AI Gateway for securing enterprise agents.

Pricing: Open-source gateway free; developer and enterprise plans via Portkey and Palo Alto Networks.

NannyML

Post-deployment ML monitoring that estimates model performance without ground-truth labels. Detects data and concept drift on tabular models and ties alerts to business impact. Acquired by Soda in 2025; open-source library still maintained.

Pricing: Open source; commercial license and cloud; see site.

Laminar

Open-source observability platform for AI agents with tracing, evaluation, and debugging workflows. Helps teams inspect long-running agent behavior, diagnose failures, and improve reliability in production deployments. Well suited for engineering teams operationalizing agentic systems at scale.

Pricing: Open-source core available. Cloud and team plans available; see site for current pricing.

Voker

Analytics platform for production AI agents. Classifies user intents, corrections, and resolutions from multi-turn conversations; surfaces queryable timelines, performance trends, and business-impact views for PMs and analysts. Lightweight Python and TypeScript SDK integrates with OpenAI, Anthropic, Gemini, LangChain, CrewAI, and Vercel AI SDK. Complements trace tools like Langfuse and LangSmith with product-facing agent insights.

Pricing: Free: 2,000 events/month. Starter: $80/month. Agent First: $400/month. Enterprise: custom; self-host and SSO on Scale.

Latitude

Open-source AI agent observability and prompt ops platform. Version prompts, run evals, and monitor LLM/agent behavior in production with Prompt Manager and PromptL.

Pricing: Open-source core; cloud plans on latitude.so.

HoneyHive

Observability layer for production AI agents: tracing, evaluation, and debugging across multi-step agent workflows. Helps teams inspect tool calls, failures, and quality regressions before and after deploy.

Pricing: Free tier and team plans; see honeyhive.ai.

TruLens

Open-source evaluation and tracing for AI apps and agents (Snowflake). Feedback functions, groundedness and relevance checks, and instrumentation for LLM and RAG pipelines. Complements production monitors with systematic evals.

Pricing: Open source (Apache 2.0); Snowflake ecosystem support.

Phoenix

Arize open-source LLM observability and eval toolkit (Apache 2.0). OpenTelemetry-native tracing, embedding analysis, and evaluation primitives for RAG and agents; self-host or use with Arize cloud.

Pricing: Open source free; Arize cloud enterprise; see docs.

Weave

Weights & Biases Weave for LLM and agent observability: traces, evaluations, and production debugging tied to the W&B experiment stack. Built for teams already tracking models in Weights & Biases.

Pricing: Included with W&B plans; see wandb.ai/site/weave.

Traceloop

LLM reliability platform built on OpenLLMetry, the open-source OpenTelemetry instrumentation for LLM apps. Turns evals and monitors into a release feedback loop. Being acquired by ServiceNow to power observability in its AI Control Tower.

Pricing: Open source and cloud; see traceloop.com.

AgentOps

Observability and session replay for AI agents across many frameworks. Time-travel debugging of tool calls, failures, and multi-step agent runs for teams shipping agents to production.

Pricing: Free tier and paid plans; see agentops.ai.

LangWatch

LLM observability and evaluation platform for traces, user analytics, and quality monitoring. Helps product teams catch regressions and improve prompts with production feedback loops.

Pricing: Plans on langwatch.ai.

Confident

Enterprise AI quality platform for standardizing LLM evaluation across product teams. Datasets, metrics, and regression workflows for reliable GenAI releases at org scale.

Pricing: Enterprise and team plans; see confident-ai.com.

DeepEval

Open-source LLM evaluation framework with unit-test style metrics for RAG, agents, and chatbots. CI-friendly scorers for hallucination, relevance, and task success.

Pricing: Open source; cloud options on deepeval.com.

Athina

Collaborative LLM eval and observability: traces, online evaluations, prompt management, and datasets. Self-host or cloud with 50+ preset evaluators. Used by teams shipping production LLM features.

Pricing: Free Starter (10k logs/month); Pro and enterprise self-host; see site.

Respan

LLM gateway plus observability (Keywords AI rebrand). Route to 1,000+ models with fallbacks, traces, cost metrics, and production evals. Used at large token volume by agent and product teams.

Pricing: Free tier; team and enterprise; see respan.ai.

Opik

Open-source AI observability and evaluation platform from Comet. Trace LLM and agent runs, run online and offline evals, and debug production GenAI systems with self-host or Comet Cloud options.

Pricing: Open source; Comet Cloud plans for hosted Opik.

Ragas

Open-source framework for evaluating RAG and LLM applications. Metrics for faithfulness, relevance, and context quality, plus synthetic eval data and production monitoring.

Pricing: Open source; enterprise via Vibrant Labs.

Lunary

Open-source observability for LLM and agent applications. Tracing, analytics, prompt management, and evaluation for teams shipping AI products.

Pricing: Open source and cloud; see Lunary.

OpenLIT

Open-source OpenTelemetry-native platform for LLM and agent observability. Auto-instrumentation, GPU metrics, and self-hosted traces for AI engineering stacks.

Pricing: Open source; see OpenLIT.

Judgment Labs

Continuous-improvement stack for AI agents. Monitors agent behavior at scale and turns production signals into evaluation and training loops. Raised $32M led by Lightspeed.

Pricing: See Judgment Labs for product and pricing.

Raindrop

Monitoring platform for production AI agents. Reads full agent trajectories to catch silent failures-hallucinated answers, tool misuse, loops, and regressions after model or prompt changes-then traces, triages in Slack or MCP, and replays traffic in Simulations before shipping. Used by Vercel, Framer, Clay, and Fortune 100 companies; $50M raised, Series A led by CRV.

Pricing: Hobby free (1,000 events/month). Pro $299/month plus usage. Enterprise custom.