Agent skill

skill-isolation-tester

Use PROACTIVELY when validating Claude Code skills before sharing or public release. Automated testing framework using multiple isolation environments (git worktree, Docker containers, VMs) to catch environment-specific bugs, hidden dependencies, and cleanup issues. Includes production-ready test templates and risk-based mode auto-detection. Not for functional testing of skill logic or non-skill code.

Stars 4
Forks 2

Install this agent skill to your Project

npx add-skill https://github.com/cskiro/claudex/tree/main/plugins/skill-isolation-tester/skills/skill-isolation-tester

SKILL.md

Skill Isolation Tester

Tests Claude Code skills in isolated environments to ensure they work correctly without dependencies on your local setup.

When to Use

Trigger Phrases:

  • "test skill [name] in isolation"
  • "validate skill [name] in clean environment"
  • "test my new skill in worktree/docker/vm"
  • "check if skill [name] has hidden dependencies"

Use Cases:

  • Test before committing or sharing publicly
  • Validate no hidden dependencies on local environment
  • Verify cleanup behavior (no leftover files/processes)
  • Catch environment-specific bugs

Quick Decision Matrix

Request Mode Isolation Level
"test in worktree" Git Worktree Fast, lightweight
"test in docker" Docker Full OS isolation
"test in vm" VM Complete isolation
"test skill X" (unspecified) Auto-detect Based on skill risk

Risk-Based Auto-Detection

Risk Level Criteria Recommended Mode
Low Read-only, no system commands Git Worktree
Medium File creation, bash commands Docker
High System config changes, VM ops VM

Mode 1: Git Worktree (Fast)

Best for: Low-risk skills, quick iteration

Process:

  1. Create isolated git worktree
  2. Install Claude Code
  3. Copy skill and run tests
  4. Cleanup

Workflow: modes/mode1-git-worktree.md

Mode 2: Docker Container (Balanced)

Best for: Medium-risk skills, full OS isolation

Process:

  1. Build/pull Docker image
  2. Create container with Claude Code
  3. Run skill tests with monitoring
  4. Cleanup container and images

Workflow: modes/mode2-docker.md

Mode 3: VM (Safest)

Best for: High-risk skills, untrusted code

Process:

  1. Provision VM, take snapshot
  2. Install Claude Code
  3. Run tests with full monitoring
  4. Rollback or cleanup

Workflow: modes/mode3-vm.md

Test Templates

Production-ready templates in test-templates/:

Template Use For
docker-skill-test.sh Docker container/image skills
docker-skill-test-json.sh CI/CD with JSON/JUnit output
api-skill-test.sh HTTP/API calling skills
file-manipulation-skill-test.sh File modification skills
git-skill-test.sh Git operation skills

Usage:

bash
chmod +x test-templates/docker-skill-test.sh
./test-templates/docker-skill-test.sh my-skill-name

# CI/CD with JSON output
export JSON_ENABLED=true
./test-templates/docker-skill-test-json.sh my-skill-name

Helper Library

lib/docker-helpers.sh provides robust Docker testing utilities:

bash
source ~/.claude/skills/skill-isolation-tester/lib/docker-helpers.sh

trap cleanup_on_exit EXIT
preflight_check_docker || exit 1
safe_docker_build "Dockerfile" "skill-test:my-skill"
safe_docker_run "skill-test:my-skill" bash -c "echo 'Testing...'"

Functions: validate_shell_command, retry_docker_command, cleanup_on_exit, preflight_check_docker, safe_docker_build, safe_docker_run

Validation Checks

Execution:

  • Skill completes without errors
  • Output matches expected format
  • Execution time acceptable

Side Effects:

  • No orphaned processes
  • Temporary files cleaned up
  • No unexpected system modifications

Portability:

  • No hardcoded paths
  • All dependencies documented
  • Works in clean environment

Test Report Format

markdown
# Skill Isolation Test Report: [skill-name]
## Environment: [Git Worktree / Docker / VM]
## Status: [PASS / FAIL / WARNING]

### Execution Results
✅ Skill completed successfully

### Side Effects Detected
⚠️ 3 temporary files not cleaned up

### Dependency Analysis
📦 Required: jq, git

### Overall Grade: B (READY with minor fixes)

Reference Materials

  • modes/mode1-git-worktree.md - Fast isolation workflow
  • modes/mode2-docker.md - Container isolation workflow
  • modes/mode3-vm.md - Full VM isolation workflow
  • data/risk-assessment.md - Skill risk evaluation
  • data/side-effect-checklist.md - Side effect validation
  • templates/test-report.md - Report template
  • test-templates/README.md - Template documentation

Quick Commands

bash
# Test with auto-detection
test skill my-new-skill in isolation

# Test in specific environment
test skill my-new-skill in worktree  # Fast
test skill my-new-skill in docker    # Balanced
test skill my-new-skill in vm        # Safest

Version: 0.1.0 | Author: Connor

Expand your agent's capabilities with these related and highly-rated skills.

cskiro/claudex

e2e-testing

Use PROACTIVELY when setting up end-to-end testing, debugging UI issues, creating visual regression suites, or automating browser testing. Uses Playwright with LLM-powered visual analysis, screenshot capture, and fix recommendations. Zero-setup for React, Next.js, Vue, Node.js, and static sites. Not for unit testing, API-only testing, or mobile native apps.

4 2
Explore
cskiro/claudex

github-repo-setup

Use PROACTIVELY when user needs to create a new GitHub repository or set up a project with best practices. Automates repository creation with four modes - quick public repos (~30s), enterprise-grade with security and CI/CD (~120s), open-source community standards (~90s), and private team collaboration with governance (~90s). Not for existing repo configuration or GitHub Actions workflow debugging.

4 2
Explore
cskiro/claudex

sub-agent-creator

Use PROACTIVELY when creating specialized Claude Code sub-agents for task delegation. Automates agent creation following Anthropic's official patterns with proper frontmatter, tool configuration, and system prompts. Generates domain-specific agents, proactive auto-triggering agents, and security-sensitive agents with limited tools. Not for modifying existing agents or general prompt engineering.

4 2
Explore
cskiro/claudex

accessibility-audit

Use PROACTIVELY when user asks for accessibility review, a11y audit, WCAG compliance check, screen reader testing, keyboard navigation validation, or color contrast analysis. Audits React/TypeScript applications for WCAG 2.2 Level AA compliance with risk-based severity scoring. Includes MUI framework awareness to avoid false positives. Not for runtime accessibility testing in production, automated remediation, or non-React frameworks.

4 2
Explore
cskiro/claudex

otel-monitoring-setup

Use PROACTIVELY when setting up OpenTelemetry monitoring for Claude Code usage tracking, cost analysis, or productivity metrics. Provides local PoC mode (full Docker stack with Grafana) and enterprise mode (centralized infrastructure). Configures telemetry collection, imports dashboards, and verifies data flow. Not for non-Claude telemetry or custom metric definitions.

4 2
Explore
cskiro/claudex

json-outputs-implementer

Use PROACTIVELY when extracting structured data from text/images, classifying content, or formatting API responses with guaranteed schema compliance. Implements Anthropic's JSON outputs mode with Pydantic/Zod SDK integration. Covers schema design, validation, testing, and production optimization. Not for tool parameter validation or agentic workflows (use strict-tool-implementer instead).

4 2
Explore

Didn't find tool you were looking for?

Be as detailed as possible for better results