test: platform-scaled hot-reload waitFor ceiling (darwin 60s, others 15s) - #4
Merged
Merged
Conversation
…15s) 2026-10-04's CI storm produced 4 hot-reload timeouts in ~30 minutes, all on darwin: FSEvents delivery latency has no SLA and exceeded the 15s ceiling on loaded shared runners. Raising the ceiling is free on healthy runs — the 50ms poll returns as soon as the condition holds; the ceiling only burns when the watcher is genuinely stuck. 60s on darwin matches the ceiling test/e2e/run.mjs's poll() already uses successfully. Timeout errors now report the platform and the actual wait, so the next flake brings data instead of a guess.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
2026-10-04 的 CI 高峰期里,热加载测试在 ~30 分钟内超时 4 次,全部在 darwin 上:FSEvents 的送达延迟没有 SLA,负载中的共享 runner 上超过了 15s 上限(失败样本 duration_ms=15039,真实尾部未知)。
What
waitFor默认上限按平台区分:darwin 60s,其他平台维持 15s(Linux inotify / Windows ReadDirectoryChangesW 从未接近过 15s)poll()已在生产使用的上限对齐,有实证依据为什么不是根治
抬上限是在买尾部余量(5→15→60 的追尾巴没有尽头)。真正确定性的方案是测试断言改轮询式 watch,但那需要改动插件的选项面,不值得为一个 CI flake 做。如果 60s 之后还出现 darwin 超时,再开 issue 讨论结构性方案。
Cost analysis
健康运行时成本为零:50ms 轮询在条件满足时立即返回,上限只在 watcher 真正卡死时才会被全额消耗(那种情况本来就该失败)。