fix: two Windows CI races - the fake servers read whole requests, and a stale bridge lock is taken over once - #36
Merged
Conversation
doctor_checks_the_ollama_model failed on main's Windows CI twice in a row, the second time with the earlier fixture fix in place. That fix read a request as far as its head, but alc's /api/show is a POST, and when its body arrived in a later packet than the head the stub answered and closed with the body unread - which Windows turns into a reset, so the client saw a broken connection and doctor reported the model "not pulled". Both stubs now read the head and then as many body bytes as Content-Length says, and close the way a real server does: answer, shut the write side, and read off anything left until the client hangs up. The quota stub had the original single-read shape and gets the same treatment. A new test sends a POST's head and body as two writes with a pause between them: against the old stub it failed five times out of five with "connection aborted"; now it passes, as does the doctor test ten times out of ten. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
only_one_of_several_starters_takes_over_a_stale_lock failed on a Windows runner: two of eight starters came away holding the bridge's starting lock. A stale lock was taken over by writing a mark over it and reading the mark back 250 ms later, and a taker that stalled longer than that between judging the lock dead and writing its mark could confirm after the first taker already had. On a loaded machine that is two bridges again - the race the lock exists to stop. A takeover now happens under a lock of its own, created exclusively and held only for the few steps it takes, and whoever holds it asks again whether the starting lock is still dead before writing over it. Every other taker finds the takeover held or the starting lock alive again, and waits for the bridge. Nothing depends on timing any more, which also drops the 250 ms every start used to wait. A takeover lock left behind by a taker that died goes stale like the starting lock and is taken over too; Windows' access-denied for a name still being deleted is asked again for a moment rather than read as a permission error. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
main's Windows CI went red after 1.11.0 merged. Two separate races, both fixed here; the second is a product fix that ships in the next release.
1.
doctor_checks_the_ollama_model(test fixture). alc's/api/showis a POST. The fake Ollama server read a request only as far as its head, so when the JSON body arrived in a later packet the stub answered and closed with the body unread. Windows turns that close into a reset, the client saw a broken connection, and doctor reported the model "not pulled". Both fake servers (serve_ollama,serve_quota) now read the head and then theContent-Lengthbody, and close gracefully: answer, shut the write side, drain until the client hangs up.2.
only_one_of_several_starters_takes_over_a_stale_lock(bridge start lock). Two of eight racing starters came away holding the bridge's starting lock. A stale lock was taken over by writing a mark over it and confirming 250 ms later, and a taker that stalled longer than that could confirm after the first one had. A takeover now happens under a short-lived lock of its own, created exclusively, and whoever holds it re-checks that the starting lock is still dead before writing over it. No timing assumption remains, and every bridge start saves the old 250 ms settle.Test plan
the_fake_ollama_server_answers_a_body_sent_after_its_head(head and body as two writes with a pause): against the old stub it failed 5 of 5 (ConnectionAborted, os error 10053); with the fix it and the doctor test passed 10 of 10cargo test --all-targets --all-featureson Windows,cargo clippy --all-targets --all-features -- -D warnings(Windows and Linux target),cargo fmt -- --check🤖 Generated with Claude Code