| title | Goals |
|---|---|
| description | Write test steps as goals with agent.act, and structure the tests around them. |
Use agent calls for goals and questions about the screen. Use locators when you know the exact action or value to check. A test can freely mix both.
import { test, expect } from 'e2e';
test('todos survive a filter round-trip', async ({ app, agent, screen }) => {
await app.open('/todos');
await agent.act('add todos "Buy milk" and "Walk the dog", complete the first');
await agent.assert('only one todo remains open');
await screen.getByRole('tab', 'Done').tap();
await expect(screen.getByTestId('todo')).toHaveCount(1);
});Give act one goal. The agent chooses and performs the actions, then returns
when the goal passes or throws when it fails. Await every agent and screen
call, and every expect.poll; a test that returns while one is still running
fails with STEP_NOT_AWAITED.
await agent.act('complete checkout with the test card');
await agent.act('invite {email} as an editor', { params: { email: '[email protected]' } });Pass values through params so the agent uses your test data. For passwords,
use a Secret, this prevents the model from seeing your secrets. See Signing in.
import { credentials } from 'e2e';
const member = credentials.user('member');
await agent.act('Sign in', {
params: { username: member.username, password: member.password },
});Name what is visible and prefer the exact wording on screen. One goal per call. The flow inside a goal is the model's job; the order of goals is yours.
For a value that changes on every run, wrap it in unique() so the
replay cache can still replay the step:
import { expect, unique } from 'e2e';
const email = `ada+${Date.now()}@example.test`;
await agent.act('sign up with email {email}', { params: { email: unique(email) } });
await expect(screen.getByText(email)).toBeVisible();If a goal fails, check the report and make the instruction more specific.
Use the labels on screen. Put vocabulary shared by several goals in the
agent's context; see Agents and personas.
A step can pass, fail, or be blocked by something such as missing credentials. How agent steps work explains the verdicts and exit codes.
act requests screenshots itself when it needs them, so unlike the
judgments it has no vision option. On
supported engines, it can tap and type at points in the image:
await agent.act('enter the code 3141 on the drawn keypad and press OK');Images add input tokens. Screenshots are withheld for the rest of an attempt after a secret is filled. The agent reference lists the modes and restrictions.
To check what a goal left on screen, follow it with an assertion or a locator check.
A test that takes only app calls no model and drives no screen. Use it to
check your API with fetch in the same run as UI tests. See
Testing APIs.
import { beforeEach, describe, test } from 'e2e';
describe('billing', { tags: ['billing'] }, () => {
beforeEach(async ({ app }) => { await app.open('/billing'); });
test('upgrades', { retries: 2, timeout: 60_000 }, async ({ agent }) => {
await agent.act('upgrade to Pro');
});
});
test.skip('not ready yet', async () => {});
test.only('just this one while I work', async () => {}); // rejected in CI
test('needs two organizations', async ({ app }) => {
test.skip((await countOrganizations()) < 2, 'one organization: nothing to switch to');
await app.open('/organizations');
});test.skip(condition, reason) inside a body throws when the condition holds:
the test stops at that line and is reported skipped, as in Playwright. Call
it before the first step when it guards a precondition.
describe and the hooks are also test.describe, test.beforeEach, and so
on. A beforeEach imported from e2e sees the core fixtures; import
test, describe, and the hooks from @e2e-dev/web (or @e2e-dev/mobile)
to have browser (or device) typed from one import, and use
test.beforeEach on a test you extended with your own fixtures.
Common options include timeout, retries, and tags. Use session for
state from a setup test, agentContext for extra instructions, and agent
to choose a configured agent. trace and video record one test
without recording the suite; see Watch it happen.
Group options apply to their tests. The test reference
lists all options and defaults.
Each test runs in its own attempt. Prepare the app data it needs through
fixtures or hooks. When a flow must span several tests, mark the group
{ serial: true }. Members share app state, run in order on one worker,
and retry as a whole.
Share setup and cleanup through test.extend. A fixture can create a test
workspace before the test and delete it afterward, including when the test
fails. See Your own fixtures.
pnpm exec e2e run tests/signup.e2e.ts:12 # the test declared at line 12
pnpm exec e2e run --grep checkout # tests whose title matches
pnpm exec e2e run --target web # one target, by namebunx e2e run tests/signup.e2e.ts:12 # the test declared at line 12
bunx e2e run --grep checkout # tests whose title matches
bunx e2e run --target web # one target, by nameThe CLI reference lists every flag.
Check the screen with `assert`, `waitFor`, and `extract`. Exact actions and checks with `screen` and `expect`.