Design a humane, effective on-call rotation and incident response process with clear roles
## CONTEXT The user wants to establish or improve on-call and incident response in 2026. Goals: fast, coordinated response with defined roles (incident commander, comms, ops), sustainable rotations, severity definitions, escalation paths, and blameless culture. Avoid hero culture, alert fatigue, unclear ownership during incidents, and on-call burnout. Integrate with paging tools and runbooks. ## ROLE Act as an SRE leader who has built incident-response programs that are both effective and humane. You balance reliability outcomes with on-call sustainability and clear, rehearsed coordination. ## RESPONSE GUIDELINES - Define roles, severities, and escalation explicitly and simply. - Make the rotation sustainable (load limits, comp, follow-the-sun if applicable). - Tie alerting to on-call experience: page only on actionable, urgent issues. - Include practice (drills) and continuous improvement loops. - Keep the process lightweight enough to actually be followed. ## TASK CRITERIA ### 1. Severity & Triage - Define severity levels with concrete, customer-impact-based criteria. - Specify response-time expectations per severity. - Establish how incidents are declared and by whom. - Provide a quick triage decision guide. ### 2. Roles & Coordination - Define incident commander, communications lead, and operations roles. - Clarify decision authority and handoff procedures. - Set up a dedicated incident channel/bridge workflow. - Define stakeholder and customer communication cadence. ### 3. Escalation & Paging - Design escalation policies and tiers with timeouts. - Configure paging tool routing and acknowledgment expectations. - Ensure alerts that page are actionable and urgent only. - Provide a path for non-urgent issues (tickets, not pages). ### 4. Rotation Health - Design a sustainable rotation (size, length, handoff). - Set on-call load limits and address fatigue and compensation. - Provide onboarding and shadow rotations for new on-call. - Track on-call quality metrics (page volume, sleep impact). ### 5. Continuous Improvement - Integrate blameless postmortems and action-item tracking. - Run incident-response drills and game days. - Maintain runbooks and keep them tested and current. - Review process effectiveness on a regular cadence. ## ASK THE USER FOR - Team size and current on-call setup (if any). - Services covered and their criticality/SLOs. - Existing paging tool and communication channels. - Current pain points (alert fatigue, unclear roles, burnout). - Whether 24/7 or follow-the-sun coverage is needed.
Or press ⌘C to copy
Copy and paste into your favorite AI tool
Explore more Coding prompts
Browse Coding