All Documents
3,528 documents available
Phase 5.2 — Resilience & Governance
Plans post-launch hardening for webhook retries, queue overload protection, audit trails, and incident runbooks in a Laravel e-commerce system.
Markspace Coordination Protocol: Specification
Specifies a stigmergy-based coordination protocol for autonomous agent fleets using five mark types, decay, trust, and conflict resolution.
Monitoring & Observability
Prompts an SRE to design a full observability stack covering logging, metrics, tracing, alerting, SLOs, and incident response.
Task 18 Part 4: On-Call Playbooks and Incident Response
Defines on-call rotation schedules, escalation policies, and three incident response playbooks for critical, database, and security events.
Monitoring & Alerting Design Checklist
Guides you through defining SLOs, selecting metrics, designing alerts, building dashboards, and instrumenting code for production observability.
Phase 2B.1: Monitoring & Alerting Blueprint
Defines a three-tier monitoring hierarchy, dashboard layouts, alert configurations, and an incident response runbook for a blockchain relayer.
Monitoring Guide - HwpBridge
Defines a full monitoring stack for an HWP conversion service with Prometheus metrics, Loki logging, Grafana dashboards, alerting, tracing, and SLOs.
🔍 SafeWallet: Laporan Analisis Komprehensif
Audits a fintech platform's security, scalability, and architecture across monolith and microservices versions, scoring each area and listing prioritized fixes.
Enterprise Workflows
Guides integrating Bosun into enterprise development workflows with CI/CD pipelines, PR reviews, custom skills, and compliance reporting.
quickstart_guide
Walks new users through setting up on-call teams, escalation policies, services, schedules, and incident response in Squadcast.
Markspace Critical Review
Critiques a multi-agent coordination protocol's safety claims and provides an implementation plan for three defense-in-depth checks.
AI Workforce Playbook
Shares real-world lessons from building an AI agent team with OpenClaw, including scoring, autonomy phases, pipeline design, and failures.
Lecture 3: Development Through Specifications and AI Agents
Describes a development workflow where humans write BDD specifications and AI agents implement, test, and refactor code.
Review Validation Analysis: Codex, Gemini, and Copilot
Analyses an external AI code review, classifying each claim as valid, invalid, or already fixed for a prompt-based multi-agent system.
Integrations
Documents how to configure alert ingestion from 23+ monitoring tools and set up notification channels to responders.
Alert Escalation & Query Regression Detection
Defines severity tiers, escalation chains, slow query thresholds, and regression detection rules for a PostgreSQL-backed application.
Security Policy
Defines a vulnerability disclosure process, response commitments, and a 10-layer defence architecture for an agent security system.
First Steps
Walks through creating an admin user, team, service, escalation policy, on-call schedule, and test incident in OpsKnight.
First Steps
Guides you through creating an admin user, team, service, escalation policy, on-call schedule, and test incident in OpsKnight.
Prompt Agent Mission
Defines a prompt-crafting agent that transforms user briefs into production-ready prompts using type detection and assembly workflows.
Check if the web service is running
Defines a server health monitoring agent with alerting, dashboards, and incident response patterns for SRE teams.
For SREs — Failure Modes, SLOs, Synthetic Checks
Defines SLOs, failure modes, synthetic checks, and incident response patterns for SRE teams running production services.
Observability, Monitoring & Analytics Integrations Research
Surveys 9 observability tools' APIs, AI integrations, and proposes Compozy extension concepts for each.
Disaster Recovery - NRP: Strengthened Production Launch Roadmap
Maps a 3-4 week production launch plan for an Australian compliance platform, covering module integration, AI content generation, and alert channels.