Agent skill
research
Autonomous goal-directed iteration loop. Define goal + metric + guard, iterate with specialist agents (perf-optimizer, sw-engineer, ai-researcher) until metric improves or limit reached. Supports GPU workloads via Colab MCP and team mode for parallel strategy exploration.
Install this agent skill to your Project
npx add-skill https://github.com/majiayu000/claude-skill-registry/tree/main/skills/other/other/research-borda-home
SKILL.md
Autonomously improve a measurable metric through iterative, atomic code changes. Each iteration follows a fixed loop: Review → Ideate (specialist agent) → Modify (ONE change) → Commit (before verify) → Verify (metric) → Guard (regression check) → Decide (keep/revert) → Log (JSONL).
Unlike /optimize (single-pass profiling) or /develop (feature/fix/refactor), /research runs sustained improvement campaigns with automatic rollback and logged experiment history. It is the right tool when you want the system to autonomously explore many code changes toward a measurable goal — coverage targets, accuracy improvements, latency reductions — over many iterations.
plan <goal>— interactive wizard: scan codebase, present proposed config, write to state file<goal>— run the iteration loop directly (uses existing config or auto-detects)resume [run-id]— resume a previous run from saved state--teamflag — parallel strategy exploration: 2-3 teammates each own a different optimization axis--colabflag — route metric verification through a Colab MCP GPU runtime
MAX_ITERATIONS: 20 (ceiling: 50 — never exceed without explicit user override)
STUCK_THRESHOLD: 5 consecutive discards → escalation
GUARD_REWORK_MAX: 2 attempts before revert
VERIFY_TIMEOUT_SEC: 120 (local), 300 (--colab)
SUMMARY_INTERVAL: 10 iterations
DIMINISHING_RETURNS_WINDOW: 5 iterations < 0.5% each → warn user and suggest stopping
STATE_DIR: .claude/state/research/
Agent strategy mapping (agent_strategy in config → ideation agent to spawn):
agent_strategy |
Specialist agent | When to use |
|---|---|---|
auto |
heuristic | Default — infer from metric_cmd keywords |
perf |
perf-optimizer |
latency, throughput, memory, GPU utilization |
code |
sw-engineer |
coverage, complexity, lines, coupling |
ml |
ai-researcher |
accuracy, loss, F1, AUC, BLEU |
arch |
solution-architect |
coupling, cohesion, modularity metrics |
Auto-inference keyword heuristics (applied when agent_strategy: auto):
- metric_cmd contains
pytest,coverage,complexity→code→sw-engineer - metric_cmd contains
time,latency,bench,throughput,memory→perf→perf-optimizer - metric_cmd contains
accuracy,loss,f1,auc,train,val,eval→ml→ai-researcher
Stuck escalation sequence (at STUCK_THRESHOLD consecutive discards):
- Switch to a different agent type (rotate through:
code→ml→perf) - Spawn 2 agents in parallel with competing strategies; keep whichever improves metric
- Stop, report progress, surface to user — do not continue looping blindly
Task tracking: per CLAUDE.md, create TaskCreate entries for all known steps immediately at skill start. In plan mode, create P1/P2/P3. In default/resume mode, create steps 1-7. Mark in_progress when starting each step, completed when done. Keep statuses current — the task list is the user's live feed.
Plan Mode (Steps P1–P3)
Triggered by plan <goal>. Interactive wizard to configure a research run.
Step P1: Parse and scan
Parse <goal> from arguments. Scan the codebase to detect:
- Language and framework (Python, PyTorch, pytest, etc.)
- Available test runners or benchmark scripts
- Candidate metric commands (pytest coverage, benchmark scripts, eval scripts)
- Candidate guard commands (test suite, lint, type check)
- Files relevant to the goal (scope files)
Step P2: Present proposed config
Present the proposed config as a code block for user review. Include:
metric_cmd: [command that prints a single numeric result]
metric_direction: higher | lower
guard_cmd: [command that must pass (exit 0) on every kept commit]
max_iterations: [default 20]
agent_strategy: [auto | perf | code | ml | arch]
scope_files: [files the ideation agent may modify]
compute: local | colab
Dry-run both commands before presenting. If either fails, flag the error and propose corrections. Do not proceed to P3 until the user confirms or edits the config.
Step P3: Write config
Write the confirmed config to .claude/state/research/config.json. Print:
✓ Config saved to .claude/state/research/config.json
Run /research <goal> to start the iteration loop.
Default Mode (Steps 1–7)
Step 1: Load / build config
If .claude/state/research/config.json exists, read it. Otherwise, attempt auto-detection of metric_cmd and guard_cmd from the goal string and codebase scan (same logic as Plan P1, but non-interactive — infer reasonable defaults).
Generate a run-id = $(date +%Y%m%d-%H%M%S). Create the run directory:
.claude/state/research/<run-id>/
state.json ← iteration count, best metric, status
experiments.jsonl ← one line per iteration
Write initial state.json:
{
"run_id": "<run-id>",
"goal": "<goal>",
"config": {},
"iteration": 0,
"best_metric": null,
"best_commit": null,
"status": "running",
"started_at": "<ISO timestamp>"
}
Step 2: Precondition checks
Run all checks before touching code. Fail fast with a clear message if any fail:
- Clean git:
git status --porcelain→ must be empty. If dirty: print the dirty files and stop. - Not detached HEAD:
git rev-parse --abbrev-ref HEAD→ must not beHEAD. - Metric command produces numeric output: run
metric_cmdonce; parse stdout for a float. If no float found: show the output and stop. - Guard command passes: run
guard_cmdonce; must exit 0. If it fails: show the output and stop. --colabcheck (if flag present): verify Colab MCP tools are available by checking formcp__colab-mcp__runtime_execute_code. If unavailable, print setup instructions (see Colab MCP section) and stop.
Step 3: Select ideation agent
Apply the agent_strategy mapping from <constants>. If auto, apply keyword heuristics to metric_cmd. Log selected agent to state.json.
Step 4: Establish baseline (iteration 0)
Run metric_cmd and guard_cmd. Parse the metric value. Append to experiments.jsonl:
{
"iteration": 0,
"commit": "<HEAD sha>",
"metric": 0.0,
"delta": 0.0,
"guard": "pass",
"status": "baseline",
"description": "baseline",
"agent": null,
"confidence": null,
"timestamp": "<ISO>",
"files": []
}
Update state.json: best_metric = <baseline>, best_commit = <HEAD sha>.
Print: Baseline: <metric_cmd key> = <value>. Then proceed to Step 5.
Step 5: Iteration loop
For each iteration i from 1 to max_iterations:
Phase 1 — Review
Build context for the ideation agent:
git log --oneline -10(recent commits)- Last 10 lines of
experiments.jsonl(prior experiment results) git diff --stat HEAD~5 HEAD(scope of recent changes)
Summarize into a compact context block: goal, current metric vs baseline, delta trend, recently modified files, previous agent actions and outcomes.
Phase 2 — Ideate
Spawn the selected specialist agent with this prompt (adapt as needed):
Goal: <goal>
Current metric: <metric_cmd key> = <current value> (baseline: <baseline>, direction: <higher|lower>)
Experiment history (last 10):
<jsonl summary>
Scope files (read and modify only these): <scope_files>
Read the scope files. Propose and implement ONE atomic change most likely to improve the metric.
The change must not break <guard_cmd>.
Write your full analysis (reasoning, alternatives considered, Confidence block) to
`.claude/state/research/<run-id>/ideation-<i>.md` using the Write tool.
Return ONLY the JSON result line — nothing else after it:
{"description":"...","files_modified":[...],"confidence":0.N}
For --colab runs: the ideation agent (especially ai-researcher) may call mcp__colab-mcp__runtime_execute_code during this phase to prototype GPU code before committing.
If the Agent tool is unavailable (nested subagent context), implement the change inline and construct the JSON result manually.
Phase 3 — Verify files changed
git diff --stat. If no files changed (no-op): append to JSONL with status: no-op, skip to Phase 8 (log), continue loop.
Phase 4 — Commit
Stage only the modified files (never git add -A):
git add <files_modified from agent JSON>
git commit -m "experiment(research/i<N>): <description>"
If pre-commit hooks fail:
- Delegate to
linting-expertagent: provide the failing hook output and the modified files; ask it to fix the issues. Max 2 attempts. - If still failing after 2 attempts:
git restore --staged .+git checkout -- .to clean up, appendstatus: hook-blocked, continue loop.
Phase 5 — Verify metric
Run metric_cmd with timeout:
timeout <VERIFY_TIMEOUT_SEC> <metric_cmd>
For --colab: route through mcp__colab-mcp__runtime_execute_code instead of local Bash. Parse numeric result from output.
If timeout expires: append status: timeout, revert via git revert HEAD --no-edit, continue loop.
Phase 6 — Guard
Run guard_cmd (exit-code check only). Record pass or fail.
Phase 7 — Decide
| Condition | Action |
|---|---|
| metric improved AND guard pass | Keep commit. Update state.json: best_metric, best_commit. |
| metric improved AND guard fail | Rework: re-spawn agent with guard failure output. Max GUARD_REWORK_MAX (2) attempts. If still failing: revert. |
| metric improved AND gain < 0.1% AND change > 50 lines | Discard (simplicity override): git revert HEAD --no-edit. |
| no improvement | Revert: git revert HEAD --no-edit. |
git revert HEAD --no-edit — never git reset --hard (preserves history, not in deny list).
Phase 8 — Log
Append one JSONL record to experiments.jsonl:
{
"iteration": 1,
"commit": "<sha of experiment commit or revert>",
"metric": 0.0,
"delta": 0.0,
"guard": "pass|fail",
"status": "kept|reverted|rework|no-op|hook-blocked|timeout",
"description": "<agent description>",
"agent": "<agent type>",
"confidence": 0.0,
"timestamp": "<ISO>",
"files": []
}
Update state.json: iteration = i, status = running.
Phase 9 — Progress checks
- Summary every SUMMARY_INTERVAL iterations: print compact table (iteration, metric, delta, status) for the last N iterations.
- Stuck detection: if last
STUCK_THRESHOLDentries all havestatus: reverted|no-op|hook-blocked, trigger escalation (see<constants>). Log escalation action. - Diminishing returns: if last
DIMINISHING_RETURNS_WINDOWkept entries each improved < 0.5%, print a warning and suggest stopping. Do not auto-stop — let the user decide. - Early stop: if the goal specifies a numeric target (e.g., "achieve 90% coverage") and the metric crosses it, stop and mark
state.jsonstatus: goal-achieved.
Step 6: Results report
Write full report to tasks/output-research-<YYYY-MM-DD>.md using the Write tool. Do not print the full report to terminal.
Report structure:
## Research Run: <goal>
**Run ID**: <run-id>
**Date**: <date>
**Iterations**: <total> (<kept> kept, <reverted> reverted, <other> other)
**Baseline**: <metric> = <baseline value>
**Best**: <metric> = <best value> (<delta>% improvement)
**Best commit**: <sha>
### Experiment History
| # | Metric | Delta | Status | Description | Agent | Confidence |
|---|--------|-------|--------|-------------|-------|------------|
| ... |
### Summary
[2-3 sentences on what strategies worked, what didn't, what to try next]
### Recommended Follow-ups
- [next action]
Print compact terminal summary:
---
Research — <goal>
Iterations: <total> Kept: <kept> Reverted: <reverted>
Baseline: <metric_key> = <baseline>
Best: <metric_key> = <best> (<delta>% improvement, commit <sha>)
Agent: <agent type used>
→ saved to tasks/output-research-<date>.md
---
Update state.json: status = completed.
Step 7: Codex delegation (optional)
After confirming results, inspect applied changes (git diff <baseline_commit>...<best_commit> --stat) and identify tasks Codex can complete (inline comments on non-obvious changes, docstring updates for modified functions, test coverage for the modified path). Read .claude/skills/_shared/codex-delegation.md and apply the criteria defined there.
Resume Mode
Triggered by resume [run-id]. If no run-id given, list available runs from .claude/state/research/ and resume the most recent status: running one.
- Read
state.jsonfrom the run dir. - Validate git HEAD: if the current HEAD has diverged from
state.json.best_commitin an unexpected direction, warn and ask before continuing. - Continue the iteration loop from
state.json.iteration + 1.
Team Mode (--team)
When to trigger: goal spans multiple optimization axes (e.g., "improve training speed" = model architecture + data pipeline + compute efficiency), OR user explicitly passes --team.
Workflow:
- Lead completes Steps 1–4 (config, preconditions, baseline) solo.
- Lead identifies 2–3 distinct optimization axes from the goal + codebase analysis.
- Lead spawns 2–3 teammates (reasoning agents at
opusper CLAUDE.md §Agent Teams), each assigned a different axis and a matching ideation agent type. Each teammate runs in an isolated worktree (isolation: worktree).
Example axis assignment for "reduce training time":
- teammate-A =
ai-researcheraxis: model architecture changes - teammate-B =
perf-optimizeraxis: data pipeline and GPU utilization - teammate-C =
sw-engineeraxis: code-level optimizations (batching, caching)
Each teammate's spawn prompt must include:
Read .claude/TEAM_PROTOCOL.md and use AgentSpeak v2.
You are a research teammate. Your axis: <axis description>.
Ideation agent: <agent type>.
Run 3–5 independent iterations of the Review→Ideate→Modify→Commit→Verify→Guard→Log loop.
Baseline metric: <metric_cmd key> = <baseline>. Direction: <higher|lower>.
Scope files: <scope_files>.
Report: {axis, iterations_run, kept, best_metric, best_commit, description}
Call TaskUpdate(in_progress) when starting; TaskUpdate(completed) when done.
- Each teammate runs their iterations independently and reports results.
- Lead cherry-picks the winning commits from each axis into the main branch, tests for compatibility, runs
guard_cmd. - Lead measures combined metric, resolves conflicts if needed, writes the Step 6 report with per-axis breakdown and combined result.
- Shutdown teammates.
Note on CLAUDE.md §8: team mode uses in-process teammates that send TeammateIdle notifications on completion — the file-activity polling protocol does not apply; TeammateIdle is the liveness signal.
Colab MCP Integration (--colab)
Purpose: route metric verification and GPU code testing to a Colab notebook runtime instead of local execution. Essential for ML training metrics, CUDA benchmarks, and any workload requiring a GPU.
Setup (user must complete before running --colab):
- Add
"colab-mcp"toenabledMcpjsonServersinsettings.local.json:json{ "enabledMcpjsonServers": [ "colab-mcp" ] } - Ensure
colab-mcpserver is defined in.mcp.jsonundermcpServers(see project.mcp.json). - Open a Colab notebook with the runtime connected and execute the MCP connection cell.
How it works during a run:
- Step 2 (preconditions): checks for
mcp__colab-mcp__runtime_execute_codeavailability. - Phase 5 (verify metric): calls
mcp__colab-mcp__runtime_execute_codewithmetric_cmdinstead of localtimeout <cmd>. - Phase 2 (ideate):
ai-researcheragent can callmcp__colab-mcp__runtime_execute_codeto prototype GPU code before committing. VERIFY_TIMEOUT_SEC= 300 (vs 120 local) to account for network + GPU startup latency.
If Colab MCP is unavailable at Step 2, print:
⚠ Colab MCP not available. To enable:
1. Add "colab-mcp" to enabledMcpjsonServers in settings.local.json
2. Open a Colab notebook and connect the runtime
3. Execute the MCP connection cell in the notebook
Then re-run with --colab.
- Commit before verify is the foundational pattern — it enables a clean
git revert HEADif the metric does not improve. Never verify before committing. git revertovergit reset --hard— preserves experiment history, is not in the deny list.- Never
git add -A— always stage specific files returned by the agent JSON. - Never
--no-verify— if a pre-commit hook blocks, delegate tolinting-expertand fix. - Guard ≠ Verify — guard checks for regressions (tests, lint); verify checks the target metric. Both must pass to keep a commit.
- Scope files are read-only for guard/test files — the ideation agent must not modify test files or the metric/guard scripts themselves.
- JSONL over TSV — richer structured fields,
jq-parseable, no delimiter ambiguity; query withjq -c 'select(.status == "kept")' experiments.jsonl. - State persistence enables resume — if the loop crashes or times out,
resumepicks up exactly where it stopped. - Safety break: max iterations default is 20; the skill never exceeds MAX_ITERATIONS without a user override in config.
- Follow-up chains:
- Run improves metric →
/reviewfor quality validation of kept commits - Run finds the metric is plateauing →
/surveyfor SOTA comparison — maybe a fundamentally different approach is needed - Kept commits accumulate technical debt →
/develop refactorfor structural cleanup with test safety net - Run exposes performance ceiling →
/optimizefor a deeper profiling pass on the bottleneck
- Run improves metric →
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
agent-ops-spec
Manage specification documents in .agent/specs/. Use when user provides requirements, acceptance criteria, or feature descriptions that need to be tracked and validated against implementation.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-testing
Test strategy, execution, and coverage analysis. Use when designing tests, running test suites, or analyzing test results beyond baseline checks.
agent-ops-state
Maintain .agent state files. Use at session start, after meaningful steps, and before concluding: read/update constitution/memory/focus/issues/baseline consistently.
Didn't find tool you were looking for?