Doctor Droid logo

Doctor Droid

Free

Resolve incidents faster with Doctor Droid—the AI SRE agent for AIOps, root cause analysis, and automated remediation.

#AI SRE Agent#AIOps SaaS#incident investigation#root cause analysis#automated remediation#knowledge graph#incident reports#JIRA tickets#playbooks#service docs#historical alerts#infrastructure#dashboards#logs#code#chat#Slack workflow#automated Playbooks#dynamic alerting#deep integrations#MCP servers#internal tools#alert noise reduction#MTTR acceleration#DevOps#SRE teams
Starting Price
$99/mo
Type
Saas
Company
Deep Sea Tech Inc.

About Doctor Droid

Doctor Droid (backed by Y Combinator) is a self-learning AI SRE agent and AIOps platform that connects to your entire stack—cloud, code, and telemetry—to build a living knowledge graph. It automatically crawls metrics, logs, traces, cloud configs, repos, docs, and runbooks, mapping every entity across tools (e.g., a GitHub repo linked to a Datadog service, Grafana dashboard, K8s pods, and AWS resources). When an alert fires, the graph traces the blast radius in seconds, enabling faster incident response and root-cause analysis. The agent learns from every investigation, reducing resolution time (e.g., 65% faster on repeat incidents). It supports plain-English queries via Slack, automated runbook execution, dynamic alerting with anomaly detection (private beta), and smart model switching to lower costs. Pricing is credit-based (≈3 investigations per credit), with a Teams plan at $99/month for 99 credits, unlimited users, and a 14-day trial.

Key Features

AI SRE agent for incident response and root cause analysis
Knowledge graph built from incidents, tickets, playbooks, docs, and alerts
DroidAgent with deep, company‑specific context
Slack Agent for in‑workflow investigation and resolution
Automated Playbooks and runbook execution
Dynamic alerting with anomaly detection (private beta)
Real‑time incident triage: alert clustering and auto‑routing
Metrics‑connected context and cross‑tool correlation
80+ native integrations (Kubernetes, Grafana, GitHub, etc.)
MCP integrations for internal tools, databases, and APIs

Pros & Cons

Pros
  • Self-learning agent gets smarter with each incident, reducing resolution time on repeat issues (up to 65% faster).
  • Connects to existing stack via OAuth; no agents or code changes required—live in 30 minutes.
  • Knowledge graph automatically maps cross-tool relationships, providing context that no dashboard can surface.
  • Unlimited users on all paid plans; pricing is based on investigation credits, not seats.
  • Smart model switching lowers cost by using cheaper models for simple investigations.
  • Slack-native workflow enables investigation and resolution without leaving chat.
  • Comprehensive integrations (80+) and MCP servers for extending to internal tools.
  • Automated playbooks can execute remediation actions (service restarts, PR generation).
Cons
  • Pricing based on investigation credits may not suit teams with very high alert volumes; credits top up at $1 each.
  • No free tier available beyond the 14-day trial (free tier mentioned in some sources but not currently on pricing page).
  • Enterprise features (self-hosted VPC, BYOLLM, SSO/SCIM, dedicated SLA) require the costly Enterprise plan.
  • Dynamic alerting with anomaly detection is still in private beta, not generally available.
  • Agent accuracy depends on completeness of connected telemetry and knowledge sources.

Best For

SRE: Reduce alert noise and speed root cause analysis during on‑call with Slack‑first investigations.DevOps Engineer: Correlate metrics, logs, and deployments to pinpoint faulty services or changes.Incident Commander: Coordinate triage in Slack, share context, and route alerts to service owners automatically.Platform Engineer: Automate runbooks and first‑level diagnostics to resolve common production issues.Application Developer: Generate hotfix PRs when code exceptions trigger incidents and validate with linked telemetry.Observability Lead: Unify telemetry sources for connected context and smarter alert correlation.On‑call Responder: Ask plain‑English questions to investigate infrastructure, dashboards, logs, and code from anywhere.Reliability Leader: Capture learnings from postmortems into a knowledge graph that improves future guidance.Tooling/Integrations Owner: Extend coverage by connecting internal tools, databases, and APIs via MCP servers.Engineering Manager: Track incident trends, create reports, and reduce burnout by cutting noisy alerts.

Alternatives to Doctor Droid