Agent skill
rca
Root cause analysis workflows - systematic investigation of failures
Install this agent skill to your Project
npx add-skill https://github.com/majiayu000/claude-skill-registry/tree/main/skills/other/other/rca
SKILL.md
flowchart TD
FAIL([Failure]) --> RCA{"/rca"}
RCA -->|CI failure, no cluster| RCACI["rca:ci"]:::rca
RCA -->|HyperShift available| RCAHS["rca:hypershift"]:::rca
RCA -->|Kind available| RCAKIND["rca:kind"]:::rca
RCACI -->|Inconclusive| NEED{"Need cluster?"}
NEED -->|Yes| RCAHS
NEED -->|Reproduce locally| RCAKIND
RCACI --> ROOT[Root Cause Found]
RCAHS --> ROOT
RCAKIND --> ROOT
ROOT --> TDD["tdd:*"]:::tdd
classDef rca fill:#FF5722,stroke:#333,color:white
classDef tdd fill:#4CAF50,stroke:#333,color:white
Follow this diagram as the workflow.
RCA Skills
Root cause analysis workflows for systematic failure investigation.
Auto-Select Sub-Skill
When this skill is invoked, determine the right sub-skill based on context:
Step 1: Determine what's available
Check for HyperShift cluster:
ls ~/clusters/hcp/kagenti-hypershift-custom-*/auth/kubeconfig 2>/dev/null
Check for Kind cluster:
kind get clusters 2>/dev/null
Step 2: Route based on failure source and access
Where did the failure occur?
│
├─ CI pipeline (GitHub Actions) ─────────────────────────┐
│ │
│ Do you have a live cluster matching the CI env? │
│ │ │
│ ├─ HyperShift cluster available │
│ │ → Use `rca:hypershift` (deep investigation) │
│ │ │
│ ├─ Kind cluster available (for Kind CI failures) │
│ │ → Use `rca:kind` (reproduce locally) │
│ │ │
│ └─ No cluster │
│ → Use `rca:ci` (logs and artifacts only) │
│ → If inconclusive, ask user to create cluster │
│ │
├─ Local Kind cluster ──────────────────────────────────┐ │
│ → Use `rca:kind` (full local access) │ │
│ │ │
└─ HyperShift cluster ─────────────────────────────────┐│ │
→ Use `rca:hypershift` (full remote access) ││ │
││ │
After RCA is complete, switch to TDD for fix iteration: ◄──┘┘ │
- `tdd:ci` (CI-only) │
- `tdd:hypershift` (live cluster) │
- `tdd:kind` (local cluster) │
Available Skills
| Skill | Access | Auto-approve | Best for |
|---|---|---|---|
rca:ci |
CI logs/artifacts only | N/A | CI failures, no cluster |
rca:hypershift |
Full cluster access | All read ops | Deep investigation |
rca:kind |
Full local access | All ops | Kind failures, fast repro |
Concurrency limit: Only one
rca:kindsession at a time (one Kind cluster fits locally). Before routing torca:kind, runkind get clusters— if a cluster exists from another session, route torca:ciinstead or ask the user.
Related Skills
tdd:ci- Fix iteration after RCA (CI-driven)tdd:hypershift- Fix iteration with live clustertdd:kind- Fix iteration on Kindk8s:logs- Query and analyze component logsk8s:pods- Debug pod issues
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
Didn't find tool you were looking for?