Skip to content

Latest commit

 

History

History
90 lines (78 loc) · 4.19 KB

File metadata and controls

90 lines (78 loc) · 4.19 KB
title Bug bashes
description Run many explorations at once, then prove each bug with a failing test before anyone reads about it.

A bug bash runs several e2e explore sessions at once, each on one area of the app. Every finding they report is then reproduced with a test that fails for the reported reason. You get back only bugs that reproduced, each with a repro test and the steps to reach it, plus the explorer's screenshot and video where the engine captures them.

A coding agent runs the whole thing from the bug-bash topic of the skill:

Bug bash the checkout and account settings on this branch.

The agent plans the charters, fans them out, merges the results, and verifies each finding. Nothing here needs a service beyond the model your config already uses.

How it runs

The agent reads the app's routes and, for a branch, the diff, and writes five to ten charters: one-sentence goals, each on one area with one posture, such as a first-time user, numbers and copy, edge input, state across reloads, or error paths. Each posture runs as its own persona, a named agent whose system prompt sets the stance, picked per charter with --agent. It runs one e2e explore per charter, four at a time, each with its own --output directory so reports never collide, and merges the findings. It sets aside findings that the local environment, the design, or the seed data explain, and the explorer's own blind spots first: links that opened a new tab, text that only reads broken in the accessibility tree, infinite-scroll sentinels, lazy-loaded media. Each remaining issue gets a repro test that asserts the expected behavior, and counts as a bug only when that test fails with ASSERTION_FAILED on it; verifying subagents, one per area, each drive their own e2e mcp session to check locators against the live app. The report lists confirmed bugs first, then production risks the triage surfaced but could not verify locally, then rejected findings grouped by reason, then warnings.

The steps, commands, and rules the agent follows are the skill topic itself: npx e2e guide bug-bash.

Before you start

  • The app must serve several explorers at once. Start it once and set reuseExisting: true on the target's command, or declare the URL with port 0 so each run starts its own copy. reuseExisting does not apply when CI is set; use port 0 there. See Starting your app.
  • When the target's app.command brings up a stack of its own, such as a database from docker compose that it removes on exit, start the stack once and explore with a config that points at the running app and declares no command, so one explorer's exit never removes the database under the others.
  • Prefer a production build. A dev server that compiles a route on its first visit looks to an explorer like a link that does nothing.
  • Give each charter its own seeded account, and tell the explorers in the agent's context which integrations the local app has no keys for.
  • When a setup test signs in, start charters with e2e explore --session <name>: a setup that signs in without typing a password keeps screenshots available, where the engine captures them.
  • On mobile, each explorer needs its own device: declare one target per simulator or emulator.
  • Put test accounts under credentials; explorers fill passwords by name and never see them.
  • Every charter costs its model calls. The last lines of each run print tokens, cost, and duration.
  • A bash against a deployed site others use is read-only: no signups, no sign-ins, no submissions, nothing injection-shaped in URLs. A WAF block is the firewall working, not a finding.

Repro tests

A repro test fails until its bug is fixed, so keep it out of the suite that gates merges until then. Once the bug is fixed, the test passes and stays as its regression test.

What one `e2e explore` run does, its budgets, and its report. Live sessions for coding agents, several at once.