Agent skill
instrumentation-planning
Plan backend observability using RED + USE + 4 Golden Signals + JTBD
Install this agent skill to your Project
npx add-skill https://github.com/nexus-labs-automation/backend-observability/tree/main/skills/instrumentation-planning
SKILL.md
Instrumentation Planning
"What job is this telemetry helping someone accomplish?"
Every metric/span should answer a job. If you can't name the job, don't add the telemetry.
Framework Selection
| Service Type | Use |
|---|---|
| HTTP/gRPC APIs | RED (Rate, Errors, Duration) |
| Resources (pools, memory) | USE (Utilization, Saturation, Errors) |
| Comprehensive SRE | 4 Golden Signals |
Prioritization Tiers
| Tier | What | Examples |
|---|---|---|
| T0: Foundation | Must have | Service tags, HTTP middleware, error tracking, health checks |
| T1: Performance | Should have | RED per endpoint, DB tracing, external service spans, context propagation |
| T2: Resources | Should have | USE for pools, memory/GC, queue depth |
| T3: Business | Nice to have | JTBD tags, user tier, feature flags |
| T4: Resilience | Nice to have | Circuit breaker state, retry counts, timeouts |
JTBD Context
Link observability to user value:
| Job | Success | Failure | Friction |
|---|---|---|---|
| "Place order" | Order created <2s | Order error | Retries >0, duration >5s |
| "Process payment" | Payment success | Declined | Gateway timeout |
Anti-Patterns
- Measure everything → Noise, cost, no signal
- High-cardinality labels → User IDs explode storage
- 100% sampling → Unnecessary overhead
- Missing context → Can't debug without job/step
Output
Present prioritized instrumentation plan organized by tier, with specific metrics and implementation order.
References
Load based on need:
references/methodology/red-methodology.mdreferences/methodology/use-methodology.mdreferences/methodology/four-golden-signals.mdreferences/methodology/jtbd-for-backend.md
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
slo-alerting
Define SLIs, SLOs, and implement burn-rate alerting
health-checks
Implement liveness, readiness, and dependency health checks
cache-observability
Track cache hit rates, latency, and detect cache-related issues
request-tracing
Instrument HTTP/gRPC endpoints with distributed tracing and RED metrics
error-handling
Capture errors with rich context for debugging and alerting
database-observability
Instrument database queries, connection pools, and detect N+1 queries
Didn't find tool you were looking for?