All Documents
3,528 documents available
SPEC-015: Escalation & Attention Policy
Defines a four-level triage system for agent-to-human escalations, prioritising attention conservation over information routing.
Build Spec: pager-triage
Defines 8 PagerDuty incident triage tools with read-only defaults, confirmation-gated write operations, and an OpsGenie fallback.
Alerting Guide for FFmpeg RTMP
Defines 16 Prometheus alert rules for FFmpeg RTMP deployments, plus setup, notification, and incident response procedures.
GraphQL Query Examples
Provides 30+ GraphQL query, mutation, and subscription examples for an incident management API, covering common operational patterns.
Sentry Alerts Configuration Guide for Vacademy Platform
Defines 13 Sentry alert rules for payment and workflow failures, plus setup steps, channels, and response procedures.
Escalation Policy
Defines severity levels, triggers, and a notification path for incidents involving autonomous agents, with a rollback procedure for Phase B degradation.
Security Policy
Defines how to report vulnerabilities, outlines incident response SLAs, and documents a resolved credential exposure for JA4proxy.
PM Dev-Session Briefing Quality: Structured Dispatch Template
Defines a structured briefing template for PM-to-dev session dispatches, replacing a minimal 5-field prompt with six required context fields and a calibration example.
support
Documents the full lifecycle of a Hyland/Alfresco support case, from logging through escalation to closure, including severity definitions, service levels, and contact channels.
Monitoring & Alerting Design Checklist
Guides you through defining SLOs, selecting metrics, designing alerts, building dashboards, and instrumenting code for production observability.
Security & DevOps
Documents a defense-in-depth security model and DevOps pipeline for an IoT platform, covering authentication, RLS, scanning, CI/CD, and incident response.
webhooks
Explains how to configure outgoing webhooks in Squadcast to push incident events to external systems via HTTP POST, PUT, or PATCH.
Notification & Escalation Procedure
Defines a three-tier notification and escalation policy for production incidents, using PagerDuty and Slack channels.
Phase 5.2 — Resilience & Governance
Plans post-launch hardening for webhook retries, queue overload protection, audit trails, and incident runbooks in a Laravel e-commerce system.
Alert Notification Setup Guide
Walks through configuring email and webhook alert notifications for Elastic Security, including Gmail, AWS SES, SendGrid, Slack, and custom backend webhooks.
OPERATIONS RUNBOOK: AI-Driven Cultural Heritage Preservation App
Documents deployment, monitoring, incident response, and disaster recovery procedures for an AI-driven cultural heritage preservation system on AWS.
Incident Readiness Auditor
Defines a structured audit process for incident readiness, covering on-call, alerts, runbooks, postmortems, and SLOs.
Provider Backlog: Terraform Provider Hyperping
Prioritises 35 backlog items for a Terraform provider, combining competitive analysis, API audits, and undocumented field discovery.
Monitoring and Alerting Configuration
Defines Sentry error tracking, uptime monitoring with UptimeRobot and Checkly, performance thresholds, and alerting channels for a static site.
Integrations
Documents how to configure alert ingestion from 23+ monitoring tools and set up notification channels like Slack, SMS, and email.
Phase 2B.1: Monitoring & Alerting Blueprint
Defines a three-tier monitoring hierarchy, dashboard layouts, alert configurations, and an incident response runbook for a blockchain relayer.
Using Cypilot — A Real Conversation (Story)
Shows a real transcript of Cypilot mode routing requests, loading context, and gating file writes during a macOS tool build.
Governance Profile Specification
Defines a governance profile for ADL that adds compliance frameworks, autonomy tiers, and audit controls for regulated enterprise agents.
SIFEN - Monitoreo y Observabilidad
Defines a full observability stack for a SIFEN API with Prometheus, Loki, Tempo, Jaeger, and Grafana, including metrics, logs, traces, dashboards, alerts, and runbooks.