Agent skill
e2e-testing
Guide for running end-to-end tests of the Qwen Code CLI, including headless mode, MCP server testing, and API traffic inspection. Use this skill whenever you need to verify CLI behavior with real model calls, reproduce user-reported bugs end-to-end, test MCP tool integrations, or inspect raw API request/response payloads. Trigger on mentions of E2E testing, headless testing, MCP tool testing, or reproducing issues.
Install this agent skill to your Project
npx add-skill https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing
SKILL.md
E2E Testing Guide
How to run the Qwen Code CLI end-to-end — from building the bundle to inspecting raw API traffic. Use when unit tests aren't enough and you need to verify behavior through the full pipeline (model API → tool validation → tool execution).
Which binary to use
- Reproducing bugs: use the globally installed
qwencommand — this matches what the user ran when they filed the issue. - Verifying fixes: build first (
npm run build && npm run bundle), then runnode dist/cli.js— this tests your local changes.
Headless Mode
Run the CLI non-interactively with JSON output (<qwen> = qwen or
node dist/cli.js per above):
<qwen> "your prompt here" \
--approval-mode yolo \
--output-format json \
2>/dev/null
The JSON output is a stream of objects. Key types:
type: "system"— init:tools,mcp_servers,model,permission_modetype: "assistant"— model output:content[].typeistext,tool_use, orthinkingtype: "user"— tool results:content[].typeistool_resultwithis_errortype: "result"— final output withresulttext andusagestats
Pipe through jq to filter the verbose stream, e.g. extract tool-result errors:
... 2>/dev/null | jq 'select(.type=="user") | .message.content[] | select(.is_error)'
Inspecting Raw API Traffic
When debugging model behavior (wrong tool arguments, schema issues), enable API logging to see the exact request/response payloads:
<qwen> "prompt" \
--approval-mode yolo \
--output-format json \
--openai-logging \
--openai-logging-dir /tmp/api-logs
Each API call produces a JSON file (can be 80KB+ due to full message history).
The bulk is in request.messages (conversation history). Trimmed structure:
{
"request": {
"model": "coder-model",
"messages": [
{ "role": "system|user|assistant", "content": "...", "tool_calls?": [...] }
],
"tools": [
{
"type": "function",
"function": {
"name": "tool_name",
"description": "...",
"parameters": { ... } // schema sent to the model
}
}
]
},
"response": {
"choices": [
{
"message": {
"role": "assistant",
"content": "...", // text response (may be null)
"tool_calls": [
{
"id": "call_...",
"function": {
"name": "tool_name",
"arguments": "..." // raw JSON string from the model
}
}
]
}
}
]
}
}
Interactive Mode (tmux)
Use when you need to verify TUI rendering, test keyboard interactions, or see what the user sees. Headless mode is simpler when you only need structured output.
Launching
tmux new-session -d -s test -x 200 -y 50 \
"cd /tmp/test-dir && <qwen> --approval-mode yolo"
sleep 3 # wait for TUI to initialize
Sending prompts
Split text and Enter with a short delay — sending them together can cause the TUI to swallow the submit:
tmux send-keys -t test "your prompt here"
sleep 0.5
tmux send-keys -t test Enter
Waiting for completion
Poll for the input prompt to reappear instead of blind sleeping:
for i in $(seq 1 60); do
sleep 2
tmux capture-pane -t test -p | grep -q "Type your message" && break
done
Capturing output
tmux capture-pane -t test -p -S -100 # -S -100 = 100 lines of scrollback
Limitations
- Key combos:
tmux send-keyscannot reliably send all key combinations.C-?,C-Shift-*, and function keys with modifiers are unsupported or unreliable. For these, use theInteractiveSessionharness inintegration-tests/interactive/or test manually. - Visual artifacts:
capture-panecaptures the final rendered frame, not intermediate states. Flicker, tearing, or brief blank frames cannot be detected this way.
Cleanup
tmux kill-session -t test
MCP Server Testing
For testing MCP tool behavior end-to-end, read references/mcp-testing.md. It
covers the setup gotchas (config location, git repo requirement) and includes
a reusable zero-dependency test server template in scripts/mcp-test-server.js.
Token Usage Stats
Use scripts/token-stats.py to summarize token usage across recent API logs:
python3 .qwen/skills/e2e-testing/scripts/token-stats.py 20 # last 20 requests
Shows input, cached, and output tokens per request with cache hit rates. Useful for verifying prompt caching behavior or investigating unexpected token counts.
Tips
- Use interactive (tmux) mode when the bug involves permission prompts, slash commands, or keyboard interactions. Headless mode has no TUI — these don't exist there.
- Use interactive (tmux) mode for hang-related issues. Headless mode produces no output when the process stalls, giving you nothing to work with.
- Use
--approval-mode defaultwhen testing permission rules.yolobypasses rule evaluation entirely — it can't test whether a rule matches.
Recommended Agent Skills
Expand your agent's capabilities with these related and highly-rated skills.
terminal-capture
Automates terminal UI screenshot testing for CLI commands. Applies when reviewing PRs that affect CLI output, testing slash commands (/about, /context, /auth, /export), generating visual documentation, or when 'terminal screenshot', 'CLI test', 'visual test', or 'terminal-capture' is mentioned.
structured-debugging
Hypothesis-driven debugging methodology for hard bugs. Use this skill whenever you're investigating non-trivial bugs, unexpected behavior, flaky tests, or tracing issues through complex systems. Activate proactively when debugging requires more than a quick glance — especially when the first attempt at a fix didn't work, when behavior seems "impossible", or when you're tempted to blame an external system (model, API, library) without evidence.
docs-audit-and-refresh
Audit the repository's docs/ content against the current codebase, find missing, incorrect, or stale documentation, and refresh the affected pages. Use when the user asks to review docs coverage, find outdated docs, compare docs with the current repo, or fix documentation drift across features, settings, tools, or integrations.
qwen-code-claw
Use Qwen Code as a Code Agent for code understanding, project generation, features, bug fixes, refactoring, and various programming tasks
docs-update-from-diff
Review local code changes with git diff and update the official docs under docs/ to match. Use when the user asks to document current uncommitted work, sync docs with local changes, update docs after a feature or refactor, or when phrases like "git diff", "local changes", "update docs", or "official docs" appear.
qc-helper
Answer any question about Qwen Code usage, features, configuration, and troubleshooting by referencing the official user documentation. Also helps users view or modify their settings.json. Invoke with `/qc-helper` followed by a question, e.g. `/qc-helper how do I configure MCP servers?` or `/qc-helper change approval mode to yolo`.
Didn't find tool you were looking for?