What the model sees
Model requests separate runner policy and trusted project context from app content. App content cannot grant tools, credentials, or a larger budget. How agent steps work explains what a step shows the model and what it may do; this page lists the trust boundary.
Secure fields appear as
value=<secure>. Known secret values in model input
and reports become <secret:name>. This redaction has limits
for transformed values, images, and app logs.
What the model may do
Built-in tools have closed schemas and are offered only when the engine supports them. Project tools must usedefineTool and cannot replace a
built-in name. Only read-only project tools may request an observation.
A call to a tool the step does not offer, or with an argument its schema does
not declare, never runs: the model reads the refusal and tries again. A node
id that is not on the current screen fails the action with
LOCATOR_NOT_FOUND. The runner never evaluates model text as code,
selectors, shell commands, or config.
A tool call that breaks a rule fails with POLICY_DENIED:
Navigation and origins
Navigation and secret fills have no origin allowlist. The agent can follow the app to other sites and use the secrets supplied to the current step there. Password fills still require a password field. The browser engine uses a hostname-based site check for two purposes:- Configured
headersare added to requests on the target’s site. - Child frames outside that site, and
data:frames, are omitted from observations.
vercel.app count as one site, so
headers can reach other hosts under that domain. basicAuth is broader:
it answers a 401 challenge from any origin.
The runner checks the scheme on explicit navigation. app.open,
browser.goto, and the agent’s navigate verb admit http:, https:,
and the exact about:blank. A device link may use an app’s custom scheme, so device.openLink
refuses a list instead: file:, data:, javascript:, view-source:,
blob:, and filesystem:. device.openApp and the open_app tool take an
app id and refuse a link, since the device would open one as a URL. The runner does not re-check redirects or attach to popup
pages. Register browser.onDialog to handle browser dialogs.
Secrets
ASecret is an opaque handle with no plaintext accessor. Three sinks
accept one: a locator’s fill, the params of an agent step, and an engine
option that declares it (web({ basicAuth: { password } })). The
Signing in guide covers usage; the rules are here.
Before a model-directed fill, the runner checks that the step declared the
secret, the secret is configured, and the target is an enabled editable
input. Passwords require a password field. It then resolves the value on the
host and passes it to the engine. Replay repeats these checks.
A deterministic fill(secret) uses the field chosen by test code without
those model-directed field checks. It still registers the value for
redaction and disables pixels for the rest of the attempt.
Screenshots after a secret fill
After any secret fill, for the rest of the attempt, screenshot and pixel tools are unavailable,app.screenshot() is denied, and the runner omits
assertion and failure screenshots. The app could display the value anywhere,
beyond the field’s masked rectangle.
A test that restores a session whose setup filled a secret keeps
this protection. A setup that signs in without filling a secret, for example
by setting a session cookie with browser.setCookies, leaves screenshots
available to the tests that restore it. A value it uses that is not a
configured secret is not redacted.
Secrets in engine options
A secret an engine option holds (web({ basicAuth: { password: secrets.get(name) } }))
is resolved on the host when each attempt starts and registered for redaction
then, together with what the engine derives from it (for basic auth, the
base64 user:password the Authorization header carries). It is not a fill
and is protected as text only: reports, screen.txt, failure pages, what the
model and an e2e mcp session read, text downloads, and the Playwright
trace’s text (the options the browser opened with, and the Authorization: Basic header of every request) are redacted, while screenshots, the trace’s
screencast frames, and the model’s pixels are kept. A page that renders the
password on screen is not masked in pixels.
Redaction
The runner replaces every occurrence of a registered secret value with<secret:name> in model input, reports, a worker process’s console output,
trace text, and text downloads. That includes a value the test passes as a
plain string rather than a secrets.get() handle: test and describe
titles, step labels (an app.open URL, a locator’s text), a fill, an
agent.act or agent.assert instruction, the strings in agent.act
params, agentContext, and an e2e explore goal. A plain string is not a
secret fill: it is protected as text only, and screenshots stay available. Titles are redacted as the test file is collected, so
the test id, --grep, the reporters, and the artifact and failure-page file
names all see <secret:name>, and two titles in one file that differ only
in a secret value are a duplicate title path. A custom executor’s
ctx.step and the built-in agent’s prompt carry the marker, so only a
handle can type the real value.
Matching covers the value as written, JSON-escaped, HTML-escaped, and
percent-encoded, in any letter case (so a CSS text-transform does not hide
it), and with its whitespace collapsed. Text cut at a length limit is masked
too when it ends with the first 8 or more characters of a value (half of a
value shorter than 16), as written or with its whitespace collapsed, the way
an engine collapses text before it cuts it. Every string of the observed
screen, names, text, values, test ids, attributes, selectors, and frame
paths alike, is redacted once, before the model, an executor’s tree, a
replay cache descriptor, or an end anchor reads it. Redaction rewrites only
what the runner reports; assertions still compare the text the page shows.
Secret values and their cut parts never enter cache entries.
A static value shorter than 6 characters is INVALID_CONFIG, and a provider
returning one fails the fill: every occurrence of so short a value would take
ordinary text with it.
Where redaction stops
- A transformed form of the value, such as its last four characters, is not the secret and passes through.
- A value encoded as a whole rather than character by character, base64 or a hash, is a different string and passes through.
- A value the app renders in another Unicode normalization form, decomposed accents where the secret has composed ones, passes through.
- Text that drops the whitespace inside the value passes through. Text that widens it passes through too, except on the observed screen: a field whose collapsed form holds the value is kept collapsed and masked.
- A value spread over several nodes, one character per cell as a PIN pad shows it, is not one string in the tree and passes through.
- A case mapping that changes length (
ßshown asSS) escapes the prefix rule for text cut at a length limit. - A title is redacted when its file is collected, before any provider runs, so a value a secret provider returns is not redacted from titles.
- When tests run in the runner’s own process, with the config or tests passed in memory by an embedding host, their console output is not redacted.
- A Playwright trace is rewritten before it is kept whenever its session
knows a secret value, filled or not: every static one, and a provider’s
once the session resolved it. Any run of 8 or more characters of a value,
its whitespace collapsed or not, becomes
<secret:name>, so unrelated text that shares 8 characters with a value is masked too. A base64 or base64url run that decodes to a value or such a fragment (a basic-auth header, a cookie, a token segment) is replaced whole. After a fill, screencast frames are dropped; otherwise they are kept. An image the app served that draws the secret survives. A trace the runner cannot rewrite is deleted and the attempt recordsTRACE_WITHHELD. - A video masks nothing and is labeled
redaction: "incomplete". It is never recorded unless you ask. A recording a hosted browser or device service keeps stays with that service; the report holds only its URL. - A download is what the app served. When the session knows a secret
value, every whole occurrence of it in a text download (
text/*or JSON by file extension) is replaced, as in reports, and the download is labeledredaction: "complete". Fragments and base64 runs are matched in traces only. Any other download, a binary, one that is not valid UTF-8 or cannot be rewritten, or one from a session that knows no secret value, is kept as served and labeledredaction: "incomplete". - The
headersa protected preview needs are recorded in the trace as request headers. Share that trace as you would the bypass secret. command.logis captured as the process writes it. The runner does not redact them. Keep secrets out of your app’s stdout.
Sessions
Saved sessions contain authentication state. They are encrypted and last for one run. A session also carries the secret values the setup test resolved from a provider before saving it, inside the same ciphertext, and whether a secret was filled on it. The test that restores the session redacts those values and keeps its pixels withheld, so a stored token the app echoes back never reaches a failure message, a report, or the model. The envelope names the secrets in the clear, authenticated; the values are never written unencrypted.
Invalid, expired, missing, or mismatched sessions are refused. See the
session errors for the corresponding codes.
Cache trust
Shared cache entries deserve the same review as test code. They can replay actions without consulting a model and are not signed. The runner validates entries before use. Invalid entries become misses. Files are named by a hash of the expected key and refused if larger than 1 MiB. The key holds the agent’s context only as a hash of its redacted text, so no secret value enters it. Entries store actions and a redacted summary; replay executes only the actions. Secret fills store a name and authorize the fill again. An entry recorded by one agent never replays for another, and one recorded on the app’s origin never replays on a page from another origin.e2e init ignores .e2e/cache/ in Git. If you choose to
commit recordings, review their actions and typed
values. CI defaults to read-only when no cache mode is configured.
Artifacts and logs
By default everything a run writes lands under.e2e/. The
output option or --output moves it, the
replay cache stays at cache.dir, and command.log is whatever
project-relative path you configure.
An
artifacts.store receives each artifact’s bytes as they are produced,
after trace and download redaction, with the redaction label the report
records, so a store can export only what the runner vouches for. A
cache.store receives every recorded entry and supplies every replayed one;
what comes back is validated again before use.
Outbound connections
Built-in features may connect to:- one telemetry request per CLI invocation to
eu.i.posthog.com, plus at most one per session thate2e mcpserves, unless opted out; see Telemetry - one request to
eu.i.posthog.compere2e feedbackyou run; see Feedback - the provider endpoint used by your configured model. Deterministic tests and fully replayed action steps make no model call.
- readiness probes against the
readyUrlof the app the runner starts - a Playwright browser download, once, when the browser is missing
- the DevTools endpoint you name in
web({ connect }) - the GitHub API, only from the opt-in
@e2e-dev/githubreporter - the Kernel API and its hosted browsers, only from
kernel() - the Expo API and the simulators’ agent-device daemons, only from
easSimulators(). It authenticates withEXPO_TOKEN, else with the eas-cli login in~/.expo/state.json: on a CI runner, setEXPO_TOKENor keep that login out of the runner’s home - the app under test, and whatever that app itself loads
--ai-trace writes a local
file without uploading it. Your test code, plugins, and app can make
additional connections.
MCP and explore
e2e mcp uses stdio and the same action authorization as a test run. It
allows four open sessions at once by default
(--max-sessions, 1 through 16) and closes idle sessions.
e2e explore offers every configured credential to every step; see
Exploring without a test.

