Skip to content

Latest commit

 

History

History
176 lines (134 loc) · 6.24 KB

File metadata and controls

176 lines (134 loc) · 6.24 KB
title Goals
description Write test steps as goals with agent.act, and structure the tests around them.

Use agent calls for goals and questions about the screen. Use locators when you know the exact action or value to check. A test can freely mix both.

import { test, expect } from 'e2e';

test('todos survive a filter round-trip', async ({ app, agent, screen }) => {
  await app.open('/todos');

  await agent.act('add todos "Buy milk" and "Walk the dog", complete the first');
  await agent.assert('only one todo remains open');

  await screen.getByRole('tab', 'Done').tap();
  await expect(screen.getByTestId('todo')).toHaveCount(1);
});

Goals: agent.act

Give act one goal. The agent chooses and performs the actions, then returns when the goal passes or throws when it fails. Await every agent and screen call, and every expect.poll; a test that returns while one is still running fails with STEP_NOT_AWAITED.

await agent.act('complete checkout with the test card');
await agent.act('invite {email} as an editor', { params: { email: '[email protected]' } });

Pass values through params so the agent uses your test data. For passwords, use a Secret, this prevents the model from seeing your secrets. See Signing in.

import { credentials } from 'e2e';

const member = credentials.user('member');
await agent.act('Sign in', {
  params: { username: member.username, password: member.password },
});
**Write instructions as you would say them out loud**

Name what is visible and prefer the exact wording on screen. One goal per call. The flow inside a goal is the model's job; the order of goals is yours.

For a value that changes on every run, wrap it in unique() so the replay cache can still replay the step:

import { expect, unique } from 'e2e';

const email = `ada+${Date.now()}@example.test`;
await agent.act('sign up with email {email}', { params: { email: unique(email) } });
await expect(screen.getByText(email)).toBeVisible();

If a goal fails, check the report and make the instruction more specific. Use the labels on screen. Put vocabulary shared by several goals in the agent's context; see Agents and personas.

A step can pass, fail, or be blocked by something such as missing credentials. How agent steps work explains the verdicts and exit codes.

act requests screenshots itself when it needs them, so unlike the judgments it has no vision option. On supported engines, it can tap and type at points in the image:

await agent.act('enter the code 3141 on the drawn keypad and press OK');

Images add input tokens. Screenshots are withheld for the rest of an attempt after a secret is filled. The agent reference lists the modes and restrictions.

To check what a goal left on screen, follow it with an assertion or a locator check.

Tests without a screen

A test that takes only app calls no model and drives no screen. Use it to check your API with fetch in the same run as UI tests. See Testing APIs.

Structure

import { beforeEach, describe, test } from 'e2e';

describe('billing', { tags: ['billing'] }, () => {
  beforeEach(async ({ app }) => { await app.open('/billing'); });

  test('upgrades', { retries: 2, timeout: 60_000 }, async ({ agent }) => {
    await agent.act('upgrade to Pro');
  });
});

test.skip('not ready yet', async () => {});
test.only('just this one while I work', async () => {}); // rejected in CI

test('needs two organizations', async ({ app }) => {
  test.skip((await countOrganizations()) < 2, 'one organization: nothing to switch to');
  await app.open('/organizations');
});

test.skip(condition, reason) inside a body throws when the condition holds: the test stops at that line and is reported skipped, as in Playwright. Call it before the first step when it guards a precondition.

describe and the hooks are also test.describe, test.beforeEach, and so on. A beforeEach imported from e2e sees the core fixtures; import test, describe, and the hooks from @e2e-dev/web (or @e2e-dev/mobile) to have browser (or device) typed from one import, and use test.beforeEach on a test you extended with your own fixtures.

Common options include timeout, retries, and tags. Use session for state from a setup test, agentContext for extra instructions, and agent to choose a configured agent. trace and video record one test without recording the suite; see Watch it happen. Group options apply to their tests. The test reference lists all options and defaults.

Each test runs in its own attempt. Prepare the app data it needs through fixtures or hooks. When a flow must span several tests, mark the group { serial: true }. Members share app state, run in order on one worker, and retry as a whole.

Share setup and cleanup through test.extend. A fixture can create a test workspace before the test and delete it afterward, including when the test fails. See Your own fixtures.

Run one test

```bash npm npx e2e run tests/signup.e2e.ts:12 # the test declared at line 12 npx e2e run --grep checkout # tests whose title matches npx e2e run --target web # one target, by name ```
pnpm exec e2e run tests/signup.e2e.ts:12   # the test declared at line 12
pnpm exec e2e run --grep checkout          # tests whose title matches
pnpm exec e2e run --target web             # one target, by name
bunx e2e run tests/signup.e2e.ts:12   # the test declared at line 12
bunx e2e run --grep checkout          # tests whose title matches
bunx e2e run --target web             # one target, by name

The CLI reference lists every flag.

Next

Check the screen with `assert`, `waitFor`, and `extract`. Exact actions and checks with `screen` and `expect`.