Conversation
Toward #412. Fixture generation is ~85% of the test job, which now takes 15m33s. The script ran `xcodebuild test` three times, and each of those builds before it runs — so the same target was compiled three times to produce three sets of results. CI timestamps show what that costs: the second invocation runs a single test (-only-testing FirstSuite/testOne) and still takes 196 seconds. That is per-invocation build and install overhead, not test execution. Now builds once with build-for-testing into a shared derivedDataPath, then runs each pass with test-without-building against it. Local timing cannot validate this and is not offered as evidence: this machine's DerivedData is already warm, so the old script pays almost no rebuild cost here (85s for all three passes, versus 790s on CI). The saving is real only on a cold runner, so CI has to measure it. Verified locally for correctness rather than speed: all three fixtures are produced, `swift test` passes 23 tests with 1 skipped, and SanityResults still contains exactly one test, so -only-testing is still honoured under test-without-building.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughThe test preparation script now builds the app once with ChangesTest Build Reuse
Estimated code review effort: 2 (Simple) | ~10 minutes Sequence Diagram(s)sequenceDiagram
participant prepareTestResults.sh
participant xcodebuild
participant SharedDerivedData
prepareTestResults.sh->>xcodebuild: build-for-testing
xcodebuild->>SharedDerivedData: store test artifacts
prepareTestResults.sh->>xcodebuild: run functional tests without building
prepareTestResults.sh->>xcodebuild: run sanity tests without building
prepareTestResults.sh->>xcodebuild: run retry tests without building
xcodebuild->>SharedDerivedData: read shared test artifacts
Possibly related issues
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
|
Closing unmerged: measured, and it is not an improvement. What the measurement showed
Why the premise was wrongThe PR assumed the three Building is 39 seconds of a ~900 second job. The second pass runs a single test under So the per-invocation cost is simulator boot, app install, and test-runner setup/teardown. What this means for #412The target is the number of |
Toward #412. The
testjob now takes 15m33s, of which 790s (85%) is fixture generation.The cost is per-invocation build overhead, not test execution
prepareTestResults.shranxcodebuild testthree times, and that action builds before it runs — so the same target was compiled three times to produce three sets of results. CI timestamps from run31372309618:RetryTestsThe middle one is the tell:
-only-testing:SampleAppUITests/FirstSuite/testOneruns one test and still costs over three minutes. That is build-and-install overhead.This builds once with
build-for-testinginto a shared-derivedDataPath, then runs each pass withtest-without-building.Local timing is not offered as evidence
This machine's DerivedData is warm, so the old script pays almost no rebuild cost here — 85s for all three passes locally, versus 790s on CI. Measured both ways; the new version is 92s locally, i.e. indistinguishable. The saving exists only on a cold runner, so CI has to be the judge. That is the point of this PR being measurable rather than argued.
Verified locally for correctness, not speed
swift test→ 23 tests, 1 skipped, 0 failuresSanityResults.xcresultstill contains exactly 1 test, 1 passed, so-only-testingis still honoured undertest-without-buildingandSanityTests' assertions holdAn unrelated finding, worth recording on #398
While checking whether this change altered fixture content,
SanityResults.xcresultcame out at 152K on one run and 51M on another — a 340× swing. I re-ran the old, unmodified script and it also produced 51M, so this change is not the cause.The difference is a ~30MB screen recording being attached or not. So the sample-suite flakiness (#398) does not merely move a test between pass and fail buckets — it changes whether large attachments exist in the fixture at all. Any future assertion about attachment presence would be flaky for the same reason.
🤖 Generated with Claude Code
Summary by CodeRabbit