
from System Design Skills21
Guidance on designing production monitoring, including metrics, logs, traces, SLOs, and health checks to prevent invisible failures.
This skill helps agents define exactly what a system should measure to be debuggable and alertable in production. It moves the design from 'it works' to 'we know when it's broken' by applying industry-standard measurement frames.
Activate this skill when the user is designing a production system and hasn't specified how to detect failures, or when they specifically ask about monitoring, Prometheus, Grafana, Datadog, or SLOs.
Primarily designed for Claude Code and agents capable of architectural reasoning and system design.
This skill has not been reviewed by our automated audit pipeline yet.