All Documents

3,528 documents available

MONITORING.md

Check if the web service is running

Defines a server health monitoring agent with alerting, dashboards, and incident response patterns for SRE teams.

aiagent
0
1
viksant
MONITORING.md

__ALERT CONFIGURATION GUIDE__

Walks through configuring email and webhook notifications for an XRPL validator monitoring dashboard with 18 pre-built alert rules.

aiautomationsafety
0
1
realgrapedrop
AUTOMATION.md

Dagster Alerting and Automation

Describes how to route Dagster quality failures into alerts, hooks, and automated remediation while keeping SLA semantics stable.

aievalautomation
0
0
seadonggyun4
MONITORING.md

🔍 SafeWallet: Laporan Analisis Komprehensif

Audits a fintech platform's security, scalability, and architecture across monolith and microservices versions, scoring each area and listing prioritized fixes.

aievalgemini
0
1
kazanaruishere-max
RUNBOOK.md

First Steps

Guides you through creating an admin user, team, service, escalation policy, on-call schedule, and test incident in OpsKnight.

aiworkflow
0
0
Dushyant-rahangdale
AGENTS.md

NovaOps v2 - AGENTS.md

Maps a multi-agent SRE incident response system that triages alerts, runs parallel analysts, validates via consensus, and escalates critical incidents through voice calls.

aiagentllm
0
1
sujeetmadihalli
CHECKLIST.md

Security Audit Report

Documents implemented security controls, identifies gaps, and provides remediation code for a Go microservices deployment on Kubernetes.

airag
0
1
everest-an
MONITORING.md

Observability, Monitoring & Analytics Integrations Research

Surveys 9 observability tools' APIs, AI integrations, and proposes Compozy extension concepts for each.

aiagentrag
0
2
compozy
MONITORING.md

awesome-sre

Curates hundreds of links to SRE talks, articles, books, and training resources from Google, Netflix, LinkedIn, and others.

ai
0
1
icopy-site
SKILL.md

AI SRE Incident Response

Defines incident classes, severity levels, Prometheus alert rules, and runbooks for AI system failures like quality regressions and cost spikes.

aillmprompt
0
2
BagelHole
AUTOMATION.md

Events API

Documents a POST endpoint for external monitoring systems to trigger, acknowledge, and resolve incidents programmatically.

ai
0
1
Dushyant-rahangdale
RUNBOOK.md

Incident Response Plan

Defines severity levels, roles, communication templates, escalation paths, and postmortem procedures for handling production incidents.

aiworkflow
0
0
sholaj
SPEC.md

Rule 19F: SRE Incident Response

Defines severity levels, escalation tiers, SLO targets, and provides Python code for incident creation, tracking, and post-incident review generation.

ai
0
0
XnimrodhunterX
PLAYBOOK.md

generated by https://github.com/hashicorp/terraform-plugin-docs

Defines the Terraform resource schema for Checkly checks, covering API, browser, and multi-step monitors with alerting and retry configuration.

ai
0
0
checkly
EVALS.md

Generate Checker Config

Guides an LLM to generate a complete config.yaml for the Checker health-monitoring system based on user infrastructure details.

aipromptclaude
0
0
imcitius
PLAYBOOK.md

generated by https://github.com/hashicorp/terraform-plugin-docs

Documents a renamed Terraform resource for TCP monitors in the Checkly provider, with schema and examples.

ai
0
0
checkly
RUNBOOK.md

how-to-integration-pagerduty

Walks through connecting Versus Incident to PagerDuty for on-call escalation with configurable acknowledgment delays.

airag
0
0
VersusControl
MONITORING.md

go-tor Alert Response Guide

Defines alert severity levels, SLIs, SLOs, and step-by-step response procedures for a Tor monitoring system.

aiworkflow
0
3
opd-ai
DEPLOYMENT.md

South Korea AI Framework Act — Compliance Mapping

Maps 15 articles of South Korea's AI Framework Act to specific Agent Governance stack modules, providing auditable compliance code examples.

aiagentrag
0
0
microsoft
RUNBOOK.md

SRO-001 On-Call & Incident Response

Defines a complete on-call schedule, incident severity levels, response procedures, and post-incident analysis workflow for SRE teams.

ai
0
2
dapi
EVALS.md

LLMTrace Implementation TODO

Tracks 100+ security and infrastructure features for an LLM proxy, each with acceptance criteria anchored to research papers.

aiagentllm
0
4
epappas
RUNBOOK.md

Incident Response & Management

Defines severity levels, escalation policies, runbook templates, automated diagnostics, and post-incident review procedures for a platform incident response system.

aiautomation
0
1
Dewscntd
MONITORING.md

Site Reliability Engineering (SRE)

Curates 200+ links to SRE culture, education, books, hiring, reliability, and tools resources from Google, Netflix, LinkedIn, and others.

ai
0
1
BINPIPE
MONITORING.md

Aquaculture Platform - SLO/SLI Definitions

Defines seven SLO/SLI targets for an aquaculture platform, including error budget calculations, burn rate alerts, and metric dependencies enforced via Prometheus.

ai
0
0
Okan-wqm
Page 44 of 147