Agent skill
benchmarking-ml-models
Runs ML model benchmarks and evaluations. Measures inference speed, memory usage, and accuracy metrics. Use for "벤치마크", "모델 평가", "성능 테스트", "inference 속도" requests.
Install this agent skill to your Project
npx add-skill https://github.com/jiunbae/agent-skills/tree/main/ml/ml-benchmark
SKILL.md
ML Benchmark
Model performance benchmarking.
Quick Benchmark
import time
import torch
# Warmup
for _ in range(10):
model(sample_input)
# Benchmark
start = time.time()
for _ in range(100):
model(sample_input)
torch.cuda.synchronize()
elapsed = time.time() - start
print(f"Avg latency: {elapsed/100*1000:.2f}ms")
Metrics
| Metric | Description | Command |
|---|---|---|
| Latency | Inference time | time.time() |
| Throughput | Samples/sec | samples / elapsed |
| Memory | VRAM usage | torch.cuda.max_memory_allocated() |
| Accuracy | Model quality | accuracy_score(y_true, y_pred) |
Benchmark Script
# Run standard benchmark
python benchmark.py --model ./model.pt --batch-size 32 --iterations 100
Output Format
## Benchmark Results: {model_name}
| Metric | Value |
|--------|-------|
| Latency (p50) | 15.2ms |
| Latency (p99) | 22.1ms |
| Throughput | 65 samples/sec |
| Memory | 4.2 GB |
| Accuracy | 92.3% |
### Configuration
- GPU: NVIDIA A100
- Batch size: 32
- Precision: FP16
Compare Models
results = {}
for model_name, model in models.items():
results[model_name] = benchmark(model)
# Generate comparison table
Best Practices
- Always warmup before measuring
- Use
torch.cuda.synchronize()for GPU - Report p50/p99 latencies
- Document hardware configuration
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
implementing-in-background
Orchestrates multiple AI agents (Claude, Codex, Gemini) for parallel implementation in the background. Separates independent tasks from planning docs, each agent writes code directly. Context-safe with auto-save. Use for "백그라운드 구현", "bg impl", "병렬 구현", "Codex로 구현", "구현해줘", "코드 작성해줘" requests.
review-fix-loop
Autonomous review-fix cycle that continuously reviews code using background-reviewer, fixes issues, and repeats until all findings are resolved. Use for "리뷰 루프", "자동 개선", "review fix loop", "리뷰 반복", "코드 개선 루프", "keep reviewing" requests.
planning-in-background
Orchestrates multiple AI agents (Claude, Codex, Gemini) for parallel planning in the background with auto-save. Agents continue running even when session hits context limits. Use for "백그라운드 기획", "bg plan", "병렬 기획", "멀티 AI 기획", "기획해줘", "N명이 기획", "계획", "플래닝", "plan", "설계" requests.
background-reviewer
Orchestrates multi-LLM parallel code review using Claude, Codex, and Gemini. Each agent reviews from a different perspective using agent personas (security, architecture, code quality, performance). Supports persona-based review via `agt persona review`. Use for "코드 리뷰", "리뷰해줘", "bg review", "멀티 리뷰", "background review", "페르소나 리뷰" requests.
managing-context
Discovers and loads relevant project context from markdown documentation before each task. Matches context documents based on keywords, file paths, and task types. Use at task start to access project plans, architecture, and implementation status.
indexing-static-context
Provides an index of global static context files in ~/.agents/. Returns appropriate static file paths for natural language queries like "내 정보", "보안 규칙". Use when other skills or agents need to locate reference information.
Didn't find tool you were looking for?