Agent skill
e2e-testing
Use PROACTIVELY when setting up end-to-end testing, debugging UI issues, creating visual regression suites, or automating browser testing. Uses Playwright with LLM-powered visual analysis, screenshot capture, and fix recommendations. Zero-setup for React, Next.js, Vue, Node.js, and static sites. Not for unit testing, API-only testing, or mobile native apps.
Install this agent skill to your Project
npx add-skill https://github.com/cskiro/claudex/tree/main/plugins/e2e-testing/skills/e2e-testing
SKILL.md
E2E Testing
Overview
This skill automates the complete Playwright e2e testing lifecycle with LLM-powered visual debugging. It detects your app type, installs Playwright, generates tests, captures screenshots, analyzes for UI bugs, and produces fix recommendations with file paths and line numbers.
Key Capabilities:
- Zero-setup automation with multi-framework support
- Visual debugging with screenshot capture and LLM analysis
- Regression testing with baseline comparison
- Actionable fix recommendations with file:line references
- CI/CD ready test suite export
When to Use This Skill
Trigger Phrases:
- "set up playwright testing for my app"
- "help me debug UI issues with screenshots"
- "create e2e tests with visual regression"
- "analyze my app's UI with screenshots"
- "generate playwright tests for [my app]"
Use Cases:
- Setting up Playwright testing from scratch
- Debugging visual/UI bugs hard to describe in text
- Creating screenshot-based regression testing
- Generating e2e test suites for new applications
- Identifying accessibility issues through visual inspection
NOT for:
- Unit testing or component testing (use Vitest/Jest)
- API-only testing without UI
- Performance/load testing
- Mobile native app testing (use Detox/Appium)
Response Style
- Automated: Execute entire workflow with minimal user intervention
- Informative: Clear progress updates at each phase
- Visual: Always capture and analyze screenshots
- Actionable: Generate specific fixes with file paths and line numbers
Quick Decision Matrix
| User Request | Action | Reference |
|---|---|---|
| "set up playwright" | Full setup workflow | workflow/phase-1-discovery.md → phase-2-setup.md |
| "debug UI issues" | Capture + Analyze | workflow/phase-4-capture.md → phase-5-analysis.md |
| "check for regressions" | Compare baselines | workflow/phase-6-regression.md |
| "generate fix recommendations" | Analyze + Generate | workflow/phase-7-fixes.md |
| "export test suite" | Package for CI/CD | workflow/phase-8-export.md |
Workflow Overview
Phase 1: Application Discovery
Detect app type, framework versions, and optimal configuration.
→ Details: workflow/phase-1-discovery.md
Phase 2: Playwright Setup
Install Playwright and generate configuration.
→ Details: workflow/phase-2-setup.md
Phase 2.5: Pre-flight Health Check
Validate app loads correctly before full test suite.
→ Details: workflow/phase-2.5-preflight.md
Phase 3: Test Generation
Create screenshot-enabled tests for critical workflows.
→ Details: workflow/phase-3-generation.md
Phase 4: Screenshot Capture
Run tests and capture visual data.
→ Details: workflow/phase-4-capture.md
Phase 5: Visual Analysis
LLM-powered analysis to identify UI bugs.
→ Details: workflow/phase-5-analysis.md
Phase 6: Regression Detection
Compare screenshots against baselines.
→ Details: workflow/phase-6-regression.md
Phase 7: Fix Generation
Map issues to source code with actionable fixes.
→ Details: workflow/phase-7-fixes.md
Phase 8: Test Suite Export
Package production-ready test suite.
→ Details: workflow/phase-8-export.md
Important Reminders
- Capture before AND after interactions - Provides context for visual debugging
- Use semantic selectors - Prefer getByRole, getByLabel over CSS selectors
- Baseline management is critical - Keep in sync with intentional UI changes
- LLM analysis is supplementary - Use alongside automated assertions
- Test critical paths first - Focus on user journeys that matter most (80/20 rule)
- Screenshots are large - Consider .gitignore for screenshots/, use CI artifacts
- Run tests in CI - Catch visual regressions before production
- Update baselines deliberately - Review diffs carefully before accepting
Limitations
- Requires Node.js >= 16
- Browser download needs ~500MB disk space
- Screenshot comparison requires consistent rendering (may vary across OS)
- LLM analysis adds ~5-10 seconds per screenshot
- Not suitable for testing behind VPNs without additional configuration
Reference Materials
| Resource | Purpose |
|---|---|
workflow/*.md |
Detailed phase instructions |
reference/troubleshooting.md |
Common issues and fixes |
reference/ci-cd-integration.md |
GitHub Actions, GitLab CI examples |
data/framework-versions.yaml |
Version compatibility database |
data/error-patterns.yaml |
Known error patterns with recovery |
templates/ |
Config and test templates |
examples/ |
Sample setups for different frameworks |
Success Criteria
- Playwright installed with browsers
- Configuration generated for app type
- Test suite created (3-5 critical journey tests)
- Screenshots captured and organized
- Visual analysis completed with issue categorization
- Regression comparison performed
- Fix recommendations generated
- Test suite exported with documentation
- All tests executable via
npm run test:e2e
Total time: ~5-8 minutes (excluding one-time Playwright install)
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
github-repo-setup
Use PROACTIVELY when user needs to create a new GitHub repository or set up a project with best practices. Automates repository creation with four modes - quick public repos (~30s), enterprise-grade with security and CI/CD (~120s), open-source community standards (~90s), and private team collaboration with governance (~90s). Not for existing repo configuration or GitHub Actions workflow debugging.
sub-agent-creator
Use PROACTIVELY when creating specialized Claude Code sub-agents for task delegation. Automates agent creation following Anthropic's official patterns with proper frontmatter, tool configuration, and system prompts. Generates domain-specific agents, proactive auto-triggering agents, and security-sensitive agents with limited tools. Not for modifying existing agents or general prompt engineering.
accessibility-audit
Use PROACTIVELY when user asks for accessibility review, a11y audit, WCAG compliance check, screen reader testing, keyboard navigation validation, or color contrast analysis. Audits React/TypeScript applications for WCAG 2.2 Level AA compliance with risk-based severity scoring. Includes MUI framework awareness to avoid false positives. Not for runtime accessibility testing in production, automated remediation, or non-React frameworks.
otel-monitoring-setup
Use PROACTIVELY when setting up OpenTelemetry monitoring for Claude Code usage tracking, cost analysis, or productivity metrics. Provides local PoC mode (full Docker stack with Grafana) and enterprise mode (centralized infrastructure). Configures telemetry collection, imports dashboards, and verifies data flow. Not for non-Claude telemetry or custom metric definitions.
json-outputs-implementer
Use PROACTIVELY when extracting structured data from text/images, classifying content, or formatting API responses with guaranteed schema compliance. Implements Anthropic's JSON outputs mode with Pydantic/Zod SDK integration. Covers schema design, validation, testing, and production optimization. Not for tool parameter validation or agentic workflows (use strict-tool-implementer instead).
mutation-testing
Use PROACTIVELY when checking if tests catch real bugs, assessing test suite quality, finding weak tests, or measuring mutation score. Validates test effectiveness beyond coverage metrics by introducing code mutations. Supports Stryker (JS/TS), PIT (Java), mutmut (Python). Not for projects without existing test suites.
Didn't find tool you were looking for?