DevOps AI tools

AI DevOps platforms integrate artificial intelligence into the software development and operations lifecycle. These tools automate and enhance development processes, deployment workflows, and system monitoring while ensuring reliable and efficient AI system operations.

28 verified AI-first sites in DevOps.

Ciroos

AI SRE teammate for enterprise environments. Multi-agent investigations reason across systems, teams, and domains to determine what happened and why, with Signal Intelligence for noise reduction and cross-domain root-cause analysis to cut MTTR and free SRE capacity.

Pricing: Enterprise; request a demo.

Komodor

Autonomous AI SRE for Kubernetes and cloud-native stacks. Klaudia agent detects, investigates, and remediates production issues. Used by Dell, Nebius, and large platform orgs; cuts MTTR and Kubernetes waste.

Pricing: Free tier available. Enterprise pricing on request.

Overcut

Agentic workflow automation control plane for software delivery operations. Connects issue trackers, source control, and CI systems to automate SDLC tasks such as triage, code review flows, documentation updates, and remediation workflows with approvals, audit trails, and sandboxed execution.

Pricing: Enterprise pricing. Managed cloud and private deployment options.

Resolve

Always-on AI agents for running software in production. Delegates on-call triage, co-investigates incidents with engineers, and runs background operational workflows across logs, metrics, traces, and deployments. MCP, API, and custom skills for enterprise integrations.

Pricing: Enterprise; contact for deployment.

NeuBird

Agentic AI for production operations and SRE. Falcon engine correlates telemetry, topology, and changes to investigate incidents, surface predictive risks, and guide or attempt remediation across hybrid cloud. Integrates with Datadog, PagerDuty, Splunk, and 50+ ops tools.

Pricing: Enterprise; contact for pricing.

HealOps

AI SRE that closes the loop from alert to pull request. Correlates logs, traces, and configs in parallel, isolates root cause, then opens a reviewed PR with minimal diff, regression test, and linked evidence trail. Integrates with Datadog, Grafana, and PagerDuty.

Pricing: See site for plans.

Overmind

AI blast-radius analysis for infrastructure changes at merge time. Simulates the live cloud environment to show what a Terraform or config change will affect before it merges, flagging risky changes in pull requests.

Pricing: Free trial. Team plans available.

Traversal

AI SRE for complex production systems. Causal search across telemetry, code, and changes for alert triage, RCA, and self-healing. Used at DigitalOcean, Amex, and PepsiCo-scale estates.

Pricing: Enterprise; contact for deployment including BYOC.

Cleric

AI SRE that investigates alerts, proposes fixes, and learns from every resolution. Read-only by default with service maps and outcome verification. Gartner Cool Vendor in AI for SRE and observability.

Pricing: Contact for pricing.

Relvy

AI on-call engineer that investigates alerts autonomously and writes auditable notebooks. Uses runbooks, telemetry, and code context. YC-backed; SOC 2 Type II with self-host options.

Pricing: See site for plans.

Causely

Causal intelligence layer for AI ops agents. Builds a semantic model of causality across services and dependencies from existing telemetry, giving agents deterministic root cause, blast radius, and remediation context so they diagnose faster, hallucinate less, and burn fewer tokens.

Pricing: Enterprise; see Causely.

Anyshift

AI SRE on a versioned infrastructure knowledge graph. Traces outages through cloud resources, Kubernetes, and git changes; flags drift and risky config before they page.

Pricing: Contact for pricing; see Anyshift.

RunWhen

Agentic SRE platform that builds safe-for-production Skills for diagnostics and remediation. Agents run skills on alerts or on a schedule across Kubernetes and cloud estates.

Pricing: Free RunWhen Local and open Skills; platform pricing on request.

Azure SRE Agent

Microsoft's autonomous SRE agent for incident diagnosis, mitigation, and cloud operations on Azure. Connects telemetry, code, and runbooks; GA after Microsoft's own internal agent fleet.

Pricing: Pay-as-you-go Azure Agent Units; see Azure pricing.

AWS DevOps Agent

AWS frontier operations agent that investigates incidents, reviews releases, and handles SRE tasks across AWS, multicloud, and on-prem. Correlates telemetry, code, and pipelines as an always-on teammate.

Pricing: Usage-based; AWS Support credits may apply. See AWS DevOps Agent pricing.

CloudThinker

AgenticOps platform for production cloud operations. AI SRE that detects, investigates, remediates, and validates recovery across observability and ITSM tools.

Pricing: Contact for pricing; see CloudThinker.

Metoro

Kubernetes-native observability and AI SRE platform using eBPF telemetry. Guardian detects production issues, investigates root causes, verifies deployments, and generates fix pull requests without application code changes.

Pricing: Free hobby tier; paid plans from $20/node/month.

Sherlocks

Autonomous AI SRE platform for incident investigation and root-cause analysis across logs, metrics, traces, deployments, and infrastructure. Runs investigations in Slack with read-only or private VPC deployment options.

Pricing: Free tier available; enterprise pricing on request.

Adps

AI-native SRE platform with specialized agents for anomaly detection, incident investigation, remediation, and Kubernetes reliability. Continuously monitors cloud and CI/CD signals to diagnose and restore production without waiting on humans.

Pricing: Contact for demo; see adps.ai.

Hyground

Self-hosted sovereign AI SRE agent for on-prem and air-gapped clusters. Investigates incidents, automates ops skills, and keeps telemetry inside your perimeter-built for European and regulated enterprises that cannot use SaaS AIOps.

Pricing: Enterprise self-hosted; contact Hyground.

K8sGPT

CNCF Sandbox AI tool that diagnoses and remediates Kubernetes issues with pluggable LLM backends, data anonymization, and MCP server integrations for Claude and other assistants. Open-source CLI and analyzers used across the Kubernetes community.

Pricing: Open source; optional commercial AI backends.

HolmesGPT

CNCF Sandbox open-source AI SRE agent (co-maintained with Robusta lineage) that investigates alerts across Kubernetes and observability data. Runs with your chosen models for transparent, self-hosted root-cause analysis.

Pricing: Open source (Apache 2.0); see HolmesGPT docs.

Aurora

Open-source (Apache 2.0) agentic incident investigation for multi-cloud and Kubernetes. LangGraph agents query infra CLIs, knowledge graphs, and runbooks to produce RCA and gated fix PRs-self-hosted or managed SaaS from Arvo AI.

Pricing: Free self-hosted; managed SaaS free to start.

Bits

Datadog's AI agents for production ops: Bits Investigation as an AI SRE for alerts and RCA, Bits Chat over telemetry, Bits Code for production-grounded fixes, plus security and custom agent builder workflows inside the Datadog platform.

Pricing: Included with Datadog plans; 14-day free trial.

Deductive

AI SRE for production debugging and root-cause analysis, acquired by Elastic in August 2026 and headed into Elastic Observability. Agents gather evidence, test hypotheses across code, telemetry, and changes, and learn from each investigation; the standalone product remains available to existing customers.

Pricing: Enterprise; contact Deductive.

Robusta

AI SRE platform that gets smarter with every incident. Builds on the HolmesGPT open-source lineage for Kubernetes and observability investigation, remediation, and continuous learning in production.

Pricing: Open-source core; cloud and enterprise plans - see Robusta.

Middleware OpsAI

AI SRE agent for auto root-cause analysis and code-level fixes across APM, RUM, logs, and Kubernetes. Detects issues in production telemetry and proposes or applies remediations with human oversight.

Pricing: Part of Middleware platform plans; see Middleware pricing.

Cast

AI Kubernetes optimization platform that uses SLO signals to take guardrailed actions in production. Autonomously rightsizes workloads, improves performance, and cuts cloud spend across clusters.

Pricing: Free tier; paid plans and enterprise pricing on Cast.