Arize
ML and LLM observability for production models and agents. Traces, drift, embeddings, and root-cause tools, plus Phoenix open-source tracing for RAG and GenAI apps.
Pricing: Free starter. Enterprise custom; see arize.com.
Model Observability platforms provide comprehensive monitoring and analysis of AI model performance and behavior. These tools enable organizations to track model health, detect issues, and ensure reliable AI operations through continuous monitoring and analysis.
37 verified AI-first sites in Model Observability.
Last reviewed .
ML and LLM observability for production models and agents. Traces, drift, embeddings, and root-cause tools, plus Phoenix open-source tracing for RAG and GenAI apps.
Pricing: Free starter. Enterprise custom; see arize.com.
Control plane for enterprise agents. Agentic observability, continuous monitoring, enforceable policy guardrails, and auditable governance. Used by Mastercard, Nielsen, US Navy, Ally, and AIG.
Pricing: Enterprise pricing. Free guardrails tier available.
Open-source AI evaluation and observability for LLM apps, agents, and predictive models. 100+ built-in metrics, LLM-as-judge evals, drift and data-quality checks, and dashboards for tests and production monitoring. Evidently Cloud adds collaboration and tracing.
Pricing: Open Source: Free. Enterprise: Custom pricing.
Observability and evaluation platform for LLMs and AI agents. Features Luna models for low-latency evals, RAG/agent/safety/security evaluations, hallucination detection, and production guardrails. Turn offline evals into real-time monitoring. Trusted by Writer, Cisco, Clearwater Analytics, MongoDB.
Pricing: Free tier. Enterprise: Custom pricing.
Agentic management platform. Sentinel is an AI gateway that observes traffic and enforces runtime guardrails; Superwise Chat is a governed workspace for business teams.
Pricing: Free Solo edition. Enterprise; contact Superwise.
Enterprise agent discovery and governance. Finds agents across endpoints and clouds, traces reasoning and tool calls, and enforces runtime policies with audit logs.
Pricing: Enterprise; contact Arthur.
AI assistant observability that analyzes Langfuse, Braintrust, or Datadog traces overnight and opens one validated PR per day. Detects failure patterns in real conversations and proposes prompt, tool, or system-instruction fixes tested against a golden dataset before merge.
Pricing: First 3 merged PRs free; pay per merged PR after.
Validation and monitoring platform for ML models and LLMs. Features automated testing, data quality checks, and model behavior analysis. Includes drift detection, performance monitoring, and custom validation suites. Offers both open-source and enterprise solutions.
Pricing: Free open-source tier. Enterprise pricing available.
Open-source LLM and agent observability and evals, now part of ClickHouse. Traces prompts, tool calls, costs, and latency; manages prompt versions, datasets, and LLM-as-judge evaluations; integrates with OpenTelemetry, LangChain, LlamaIndex, and major model providers. Self-host or cloud.
Pricing: Self-host free. Cloud Hobby free. Core $29/month. Pro $199/month. Enterprise from $2,499/month.
Open-source AI gateway and LLM observability. Proxy-based logging of requests, costs, and latency across providers with routing, caching, and fallbacks; processed 14T+ tokens for 16,000 organizations. Acquired by Mintlify in March 2026 and now in maintenance mode-live and patched, with no new features.
Pricing: Hobby free (10k requests). Pro $79/month. Team $799/month. Enterprise custom.
AI governance, testing, and observability platform that inventories every AI system in a company. Pre-deploy tests, production monitoring for drift, latency, and prompt failures, cost controls, and audit-ready governance for both ML models and LLM agents.
Pricing: Basic trial with 20k inferences/month. Enterprise custom.
AI workflow intelligence for LLM apps-trace failures, expert review loops, and calibrated judges tied to accuracy, adoption, and business outcomes. Databricks-partnered eval/ops layer rather than classic UI regression QA.
Pricing: Contact for pricing.
LangChain-native observability for LLM applications: distributed tracing over chains and tools, datasets for regression testing, online and offline evaluation, prompt iteration, and collaboration for ML/AI teams. Tight integration with LangChain and LangGraph; positions debugging and quality gates as first-class for production GenAI-not ad hoc logging.
Pricing: Developer tier; usage-based and enterprise; see site.
Evaluation and observability platform for AI products: logged traces, scorer SDKs, prompt and experiment comparison, and CI-friendly eval workflows. Emphasizes fast iteration on LLM quality with developer-centric UX and data ownership. Complements generic APM with task-specific LLM metrics and human/LLM-as-judge patterns.
Pricing: Free tier; team and enterprise; see site.
AI research lab and evaluation company. Percival debugs agent traces and flags failure patterns; its core platform runs LLM tests, safety checks, and judges; newer work builds RL environments and Digital World Models for training frontier agents.
Pricing: Contact; enterprise; see site.
AI gateway and control plane with routing to 1,600+ models, fallbacks, observability, guardrails, governance, and prompt management. Acquired by Palo Alto Networks in May 2026 and now sold as the Prisma AIRS AI Gateway for securing enterprise agents.
Pricing: Open-source gateway free; developer and enterprise plans via Portkey and Palo Alto Networks.
Post-deployment ML monitoring that estimates model performance without ground-truth labels. Detects data and concept drift on tabular models and ties alerts to business impact. Acquired by Soda in 2025; open-source library still maintained.
Pricing: Open source; commercial license and cloud; see site.
Open-source observability platform for AI agents with tracing, evaluation, and debugging workflows. Helps teams inspect long-running agent behavior, diagnose failures, and improve reliability in production deployments. Well suited for engineering teams operationalizing agentic systems at scale.
Pricing: Open-source core available. Cloud and team plans available; see site for current pricing.
Analytics platform for production AI agents. Classifies user intents, corrections, and resolutions from multi-turn conversations; surfaces queryable timelines, performance trends, and business-impact views for PMs and analysts. Lightweight Python and TypeScript SDK integrates with OpenAI, Anthropic, Gemini, LangChain, CrewAI, and Vercel AI SDK. Complements trace tools like Langfuse and LangSmith with product-facing agent insights.
Pricing: Free: 2,000 events/month. Starter: $80/month. Agent First: $400/month. Enterprise: custom; self-host and SSO on Scale.
Open-source AI agent observability and prompt ops platform. Version prompts, run evals, and monitor LLM/agent behavior in production with Prompt Manager and PromptL.
Pricing: Open-source core; cloud plans on latitude.so.
Observability layer for production AI agents: tracing, evaluation, and debugging across multi-step agent workflows. Helps teams inspect tool calls, failures, and quality regressions before and after deploy.
Pricing: Free tier and team plans; see honeyhive.ai.
Open-source evaluation and tracing for AI apps and agents (Snowflake). Feedback functions, groundedness and relevance checks, and instrumentation for LLM and RAG pipelines. Complements production monitors with systematic evals.
Pricing: Open source (Apache 2.0); Snowflake ecosystem support.
Arize open-source LLM observability and eval toolkit (Apache 2.0). OpenTelemetry-native tracing, embedding analysis, and evaluation primitives for RAG and agents; self-host or use with Arize cloud.
Pricing: Open source free; Arize cloud enterprise; see docs.
Weights & Biases Weave for LLM and agent observability: traces, evaluations, and production debugging tied to the W&B experiment stack. Built for teams already tracking models in Weights & Biases.
Pricing: Included with W&B plans; see wandb.ai/site/weave.
LLM reliability platform built on OpenLLMetry, the open-source OpenTelemetry instrumentation for LLM apps. Turns evals and monitors into a release feedback loop. Being acquired by ServiceNow to power observability in its AI Control Tower.
Pricing: Open source and cloud; see traceloop.com.
Observability and session replay for AI agents across many frameworks. Time-travel debugging of tool calls, failures, and multi-step agent runs for teams shipping agents to production.
Pricing: Free tier and paid plans; see agentops.ai.
LLM observability and evaluation platform for traces, user analytics, and quality monitoring. Helps product teams catch regressions and improve prompts with production feedback loops.
Pricing: Plans on langwatch.ai.
Enterprise AI quality platform for standardizing LLM evaluation across product teams. Datasets, metrics, and regression workflows for reliable GenAI releases at org scale.
Pricing: Enterprise and team plans; see confident-ai.com.
Open-source LLM evaluation framework with unit-test style metrics for RAG, agents, and chatbots. CI-friendly scorers for hallucination, relevance, and task success.
Pricing: Open source; cloud options on deepeval.com.
Collaborative LLM eval and observability: traces, online evaluations, prompt management, and datasets. Self-host or cloud with 50+ preset evaluators. Used by teams shipping production LLM features.
Pricing: Free Starter (10k logs/month); Pro and enterprise self-host; see site.
LLM gateway plus observability (Keywords AI rebrand). Route to 1,000+ models with fallbacks, traces, cost metrics, and production evals. Used at large token volume by agent and product teams.
Pricing: Free tier; team and enterprise; see respan.ai.
Open-source AI observability and evaluation platform from Comet. Trace LLM and agent runs, run online and offline evals, and debug production GenAI systems with self-host or Comet Cloud options.
Pricing: Open source; Comet Cloud plans for hosted Opik.
Open-source framework for evaluating RAG and LLM applications. Metrics for faithfulness, relevance, and context quality, plus synthetic eval data and production monitoring.
Pricing: Open source; enterprise via Vibrant Labs.
Open-source observability for LLM and agent applications. Tracing, analytics, prompt management, and evaluation for teams shipping AI products.
Pricing: Open source and cloud; see Lunary.
Open-source OpenTelemetry-native platform for LLM and agent observability. Auto-instrumentation, GPU metrics, and self-hosted traces for AI engineering stacks.
Pricing: Open source; see OpenLIT.
Continuous-improvement stack for AI agents. Monitors agent behavior at scale and turns production signals into evaluation and training loops. Raised $32M led by Lightspeed.
Pricing: See Judgment Labs for product and pricing.
Monitoring platform for production AI agents. Reads full agent trajectories to catch silent failures-hallucinated answers, tool misuse, loops, and regressions after model or prompt changes-then traces, triages in Slack or MCP, and replays traffic in Simulations before shipping. Used by Vercel, Framer, Clay, and Fortune 100 companies; $50M raised, Series A led by CRV.
Pricing: Hobby free (1,000 events/month). Pro $299/month plus usage. Enterprise custom.