| title | Bug bashes |
|---|---|
| description | Run many explorations at once, then prove each bug with a failing test before anyone reads about it. |
A bug bash runs several e2e explore sessions at once, each on one area
of the app. Every finding they report is then reproduced with a test that
fails for the reported reason. You get back only bugs that reproduced, each
with a repro test and the steps to reach it, plus the explorer's screenshot
and video where the engine captures them.
A coding agent runs the whole thing from the bug-bash topic of the
skill:
Bug bash the checkout and account settings on this branch.
The agent plans the charters, fans them out, merges the results, and verifies each finding. Nothing here needs a service beyond the model your config already uses.
The agent reads the app's routes and, for a branch, the diff, and writes
five to ten charters: one-sentence goals, each on one area with one
posture, such as a first-time user, numbers and copy, edge input, state
across reloads, or error paths. Each posture runs as its own persona, a
named agent whose system prompt sets the stance, picked per charter with
--agent. It runs one e2e explore per charter, four
at a time, each with its own --output directory so reports never collide,
and merges the findings. It sets aside findings that the local
environment, the design, or the seed data explain, and the explorer's own
blind spots first: links that opened a new tab, text that only reads broken
in the accessibility tree, infinite-scroll sentinels, lazy-loaded media. Each remaining issue
gets a repro test that asserts the expected behavior, and counts as a bug
only when that test fails with ASSERTION_FAILED on it; verifying
subagents, one per area, each drive their own
e2e mcp session to check locators
against the live app. The report lists confirmed bugs first, then
production risks the triage surfaced but could not verify locally, then
rejected findings grouped by reason, then warnings.
The steps, commands, and rules the agent follows are the skill topic
itself: npx e2e guide bug-bash.
- The app must serve several explorers at once. Start it once and set
reuseExisting: trueon the target'scommand, or declare the URL with port0so each run starts its own copy.reuseExistingdoes not apply whenCIis set; use port0there. See Starting your app. - When the target's
app.commandbrings up a stack of its own, such as a database fromdocker composethat it removes on exit, start the stack once and explore with a config that points at the running app and declares no command, so one explorer's exit never removes the database under the others. - Prefer a production build. A dev server that compiles a route on its first visit looks to an explorer like a link that does nothing.
- Give each charter its own seeded account, and tell the explorers in the
agent's
contextwhich integrations the local app has no keys for. - When a setup test signs in, start charters with
e2e explore --session <name>: a setup that signs in without typing a password keeps screenshots available, where the engine captures them. - On mobile, each explorer needs its own device: declare one target per simulator or emulator.
- Put test accounts under
credentials; explorers fill passwords by name and never see them. - Every charter costs its model calls. The last lines of each run print tokens, cost, and duration.
- A bash against a deployed site others use is read-only: no signups, no sign-ins, no submissions, nothing injection-shaped in URLs. A WAF block is the firewall working, not a finding.
A repro test fails until its bug is fixed, so keep it out of the suite that gates merges until then. Once the bug is fixed, the test passes and stays as its regression test.
What one `e2e explore` run does, its budgets, and its report. Live sessions for coding agents, several at once.