Skip to content

Sample UI tests flake under CI load, making fixtures nondeterministic #398

Description

@tylervick

Now that fixtures are regenerated on every CI run rather than downloaded as fixed artifacts, flakiness in the sample app's UI tests feeds directly into the generated .xcresult — and therefore into assertions made against it.

Evidence

Commit 53adfaf, same code, two runs:

53adfaf  Test    -> success
53adfaf  Codecov -> failure

The failing run produced Passed(6)/Failed(6) where the passing run produced Passed(7)/Failed(5). Fixture-generation logs show FirstSuite.testOne as both passed and failed:

Test Case '-[SampleAppUITests.FirstSuite testOne]' passed
Test Case '-[SampleAppUITests.FirstSuite testOne]' failed

Cause

Every UI sample suite (FirstSuite, SecondSuite, ThirdSuite) calls this in setUp:

continueAfterFailure = false
XCUIApplication().launch()

testOne's body is XCTAssert(true) plus an attachment — it cannot fail on its own assertion. When the simulator is slow to launch the app under CI load, launch() fails and takes the test down with it before the body runs.

Why the launch is there at all

None of the sample tests actually interact with the app's UI. They add attachments, run activities, and assert constants. The launch() call exists so the run is a genuine UI test and produces a screen recording — which is currently the project's only .mp4 attachment coverage.

So the launch cannot simply be deleted without losing real fixture coverage.

Options

  1. Make the launch resilient — wait for a known app state, or retry the launch, rather than letting a slow simulator fail the test outright.
  2. Split the suites: keep launch() in one class that genuinely needs the recording, and drop it from the classes that only attach data.
  3. Accept the flakiness and keep assertions restricted to invariants that don't depend on it (what Stop asserting simulator reliability in testResultStatusCount #397 did as an immediate fix).

Option 1 or 2 is the real fix. #397 only stopped one assertion from being affected; the underlying nondeterminism is still there and will surface again in any future test that counts per-status results.

Why this matters beyond one test

A CI that is intermittently red for environmental reasons trains everyone to dismiss red checks — which is how a genuine failure eventually gets ignored. Worth fixing before the habit sets in.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions