e2e init installs a skill for writing and running tests, plus MCP
configuration for inspecting the app. Coding agents use the CLI to run
tests and the reports to investigate failures.
The skill
The copy holds
SKILL.md and topic references; the link is the layout
npx skills add produces, so one copy serves every agent. Run init again
after upgrading e2e to refresh it.
Without init, npx skills add tester-army/e2e installs the skill from the
repository, and npx e2e guide [topic] prints it from the
installed package.
The package also ships these docs as .mdx pages in its docs/ directory
(node_modules/e2e/docs in a single-package project), so an agent can read
any page offline. A link such as /reference/cli is docs/reference/cli.mdx.
For agents that need an explicit pointer, add this line to your project
instructions. init leaves those files unchanged:
AGENTS.md
How an agent looks at the app
e2e mcp lets the coding agent open and interact with the app. Each agent,
subagents included, opens its own session, so several can drive the app at
once. Its locate tool checks a screen.* locator against the current app and returns
the test code when exactly one element matches.
init registers the server in .mcp.json for Claude Code and
.cursor/mcp.json for Cursor. To register it with Claude Code by hand:
How an agent writes a test
The skill tells the coding agent to read the config and an existing test first. It usesagent.act for one goal at a time and checks the outcome
with expect. Known fields and exact values can use screen directly:
tests/checkout.e2e.ts
- One
agent.actper goal, followed by anexpect. The assertion makes the step’s cache entry eligible for replay, and it is what fails when the agent did the wrong thing. - No sleeps. Locators poll, actions wait for their target, and
expectretries until its timeout. - Secrets through
credentials.user(name), declared in the config, never literal in a test.
How an agent runs it
How an agent reads the result
Every run that gets to its tests writes.e2e/report.json. For a report that is easier to read,
run:
.e2e/summary.md and a page for each failed or flaky test under
.e2e/failures/. Start with the failure page. It includes the failing line,
steps, recent model turns, and the screen at failure.
When reading JSON, check run.errors for run-level failures, then the last
attempt of each failed result. Its error, steps, and artifacts show
what failed and how far the test got. --reporter json also prints the JSON
report to stdout.
For model-call details, use --ai-trace and read the trace with
unbox-ai. Avoid opening the raw trace in an agent’s
context; it repeats prompts and can be several megabytes.
How an agent treats the cache
A verifiedagent.act can replay from .e2e/cache/ on later runs. Use
--no-cache to check the live agent behavior after changing a test or to
investigate a replay failure. Pair it with --ai-trace to capture every
model call in the flow.
Do not edit recordings by hand. A failed read-write run removes implicated
entries so a later verified run can replace them. A failure where no model
answered implicates nothing and removes none. See
Caching agent steps.
CLI reference
init, guide, mcp, and run, flag by flag.
MCP server
Inspect the app before writing a test.
Debugging a run
The report, the artifacts, and the agent flags.

