All Documents
3,528 documents available
Check if the web service is running
Defines a server health monitoring agent with alerting, dashboards, and incident response patterns for SRE teams.
__ALERT CONFIGURATION GUIDE__
Walks through configuring email and webhook notifications for an XRPL validator monitoring dashboard with 18 pre-built alert rules.
Dagster Alerting and Automation
Describes how to route Dagster quality failures into alerts, hooks, and automated remediation while keeping SLA semantics stable.
🔍 SafeWallet: Laporan Analisis Komprehensif
Audits a fintech platform's security, scalability, and architecture across monolith and microservices versions, scoring each area and listing prioritized fixes.
First Steps
Guides you through creating an admin user, team, service, escalation policy, on-call schedule, and test incident in OpsKnight.
NovaOps v2 - AGENTS.md
Maps a multi-agent SRE incident response system that triages alerts, runs parallel analysts, validates via consensus, and escalates critical incidents through voice calls.
Security Audit Report
Documents implemented security controls, identifies gaps, and provides remediation code for a Go microservices deployment on Kubernetes.
Observability, Monitoring & Analytics Integrations Research
Surveys 9 observability tools' APIs, AI integrations, and proposes Compozy extension concepts for each.
awesome-sre
Curates hundreds of links to SRE talks, articles, books, and training resources from Google, Netflix, LinkedIn, and others.
AI SRE Incident Response
Defines incident classes, severity levels, Prometheus alert rules, and runbooks for AI system failures like quality regressions and cost spikes.
Events API
Documents a POST endpoint for external monitoring systems to trigger, acknowledge, and resolve incidents programmatically.
Incident Response Plan
Defines severity levels, roles, communication templates, escalation paths, and postmortem procedures for handling production incidents.
Rule 19F: SRE Incident Response
Defines severity levels, escalation tiers, SLO targets, and provides Python code for incident creation, tracking, and post-incident review generation.
generated by https://github.com/hashicorp/terraform-plugin-docs
Defines the Terraform resource schema for Checkly checks, covering API, browser, and multi-step monitors with alerting and retry configuration.
Generate Checker Config
Guides an LLM to generate a complete config.yaml for the Checker health-monitoring system based on user infrastructure details.
generated by https://github.com/hashicorp/terraform-plugin-docs
Documents a renamed Terraform resource for TCP monitors in the Checkly provider, with schema and examples.
how-to-integration-pagerduty
Walks through connecting Versus Incident to PagerDuty for on-call escalation with configurable acknowledgment delays.
go-tor Alert Response Guide
Defines alert severity levels, SLIs, SLOs, and step-by-step response procedures for a Tor monitoring system.
South Korea AI Framework Act — Compliance Mapping
Maps 15 articles of South Korea's AI Framework Act to specific Agent Governance stack modules, providing auditable compliance code examples.
SRO-001 On-Call & Incident Response
Defines a complete on-call schedule, incident severity levels, response procedures, and post-incident analysis workflow for SRE teams.
LLMTrace Implementation TODO
Tracks 100+ security and infrastructure features for an LLM proxy, each with acceptance criteria anchored to research papers.
Incident Response & Management
Defines severity levels, escalation policies, runbook templates, automated diagnostics, and post-incident review procedures for a platform incident response system.
Site Reliability Engineering (SRE)
Curates 200+ links to SRE culture, education, books, hiring, reliability, and tools resources from Google, Netflix, LinkedIn, and others.
Aquaculture Platform - SLO/SLI Definitions
Defines seven SLO/SLI targets for an aquaculture platform, including error budget calculations, burn rate alerts, and metric dependencies enforced via Prometheus.