Skip to content

fix: two Windows CI races - the fake servers read whole requests, and a stale bridge lock is taken over once - #36

Merged
treeleaves30760 merged 2 commits into
mainfrom
fix/fake-servers-read-whole-requests
Sep 22, 2026
Merged

treeleaves30760 merged 2 commits into
mainfrom
fix/fake-servers-read-whole-requests

Conversation

@treeleaves30760

@treeleaves30760 treeleaves30760 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

main's Windows CI went red after 1.11.0 merged. Two separate races, both fixed here; the second is a product fix that ships in the next release.

1. doctor_checks_the_ollama_model (test fixture). alc's /api/show is a POST. The fake Ollama server read a request only as far as its head, so when the JSON body arrived in a later packet the stub answered and closed with the body unread. Windows turns that close into a reset, the client saw a broken connection, and doctor reported the model "not pulled". Both fake servers (serve_ollama, serve_quota) now read the head and then the Content-Length body, and close gracefully: answer, shut the write side, drain until the client hangs up.

2. only_one_of_several_starters_takes_over_a_stale_lock (bridge start lock). Two of eight racing starters came away holding the bridge's starting lock. A stale lock was taken over by writing a mark over it and confirming 250 ms later, and a taker that stalled longer than that could confirm after the first one had. A takeover now happens under a short-lived lock of its own, created exclusively, and whoever holds it re-checks that the starting lock is still dead before writing over it. No timing assumption remains, and every bridge start saves the old 250 ms settle.

Test plan

  • New the_fake_ollama_server_answers_a_body_sent_after_its_head (head and body as two writes with a pause): against the old stub it failed 5 of 5 (ConnectionAborted, os error 10053); with the fix it and the doctor test passed 10 of 10
  • The lock tests, 30 of 30 under synthetic CPU load; two new deterministic tests (a takeover already in hand is not joined; a takeover lock left by a dead taker is taken over too)
  • cargo test --all-targets --all-features on Windows, cargo clippy --all-targets --all-features -- -D warnings (Windows and Linux target), cargo fmt -- --check

🤖 Generated with Claude Code

treeleaves30760 and others added 2 commits September 23, 2026 01:33
doctor_checks_the_ollama_model failed on main's Windows CI twice in a row,
the second time with the earlier fixture fix in place. That fix read a
request as far as its head, but alc's /api/show is a POST, and when its body
arrived in a later packet than the head the stub answered and closed with
the body unread - which Windows turns into a reset, so the client saw a
broken connection and doctor reported the model "not pulled".

Both stubs now read the head and then as many body bytes as Content-Length
says, and close the way a real server does: answer, shut the write side, and
read off anything left until the client hangs up. The quota stub had the
original single-read shape and gets the same treatment. A new test sends a
POST's head and body as two writes with a pause between them: against the
old stub it failed five times out of five with "connection aborted"; now it
passes, as does the doctor test ten times out of ten.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
only_one_of_several_starters_takes_over_a_stale_lock failed on a Windows
runner: two of eight starters came away holding the bridge's starting lock.
A stale lock was taken over by writing a mark over it and reading the mark
back 250 ms later, and a taker that stalled longer than that between judging
the lock dead and writing its mark could confirm after the first taker
already had. On a loaded machine that is two bridges again - the race the
lock exists to stop.

A takeover now happens under a lock of its own, created exclusively and
held only for the few steps it takes, and whoever holds it asks again
whether the starting lock is still dead before writing over it. Every other
taker finds the takeover held or the starting lock alive again, and waits
for the bridge. Nothing depends on timing any more, which also drops the
250 ms every start used to wait. A takeover lock left behind by a taker that
died goes stale like the starting lock and is taken over too; Windows'
access-denied for a name still being deleted is asked again for a moment
rather than read as a permission error.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@treeleaves30760 treeleaves30760 changed the title fix(test): the fake servers read a whole request before they answer fix: two Windows CI races - the fake servers read whole requests, and a stale bridge lock is taken over once Sep 22, 2026
@treeleaves30760
treeleaves30760 merged commit a39f9e8 into main Sep 22, 2026
7 checks passed
@treeleaves30760
treeleaves30760 deleted the fix/fake-servers-read-whole-requests branch September 23, 2026 00:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant