Skip to content

test: fix the two flaky integration tests - #33

Merged
fylorn merged 1 commit into
devfrom
fix/flaky-integration-tests
Sep 24, 2026
Merged

fylorn merged 1 commit into
devfrom
fix/flaky-integration-tests

Conversation

@fylorn

@fylorn fylorn commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Summary

successful_login_decays_subnet_failure_counter (about 1 run in 100)

  • Cause: the test client's PoW grinder gave up after 10M nonces. At difficulty 21 a solve takes 2^21 ≈ 2.1M tries on average, and the count is geometric, so 10M is exceeded with probability e^-4.77 ≈ 0.85%.
  • Fix: the cap is now 32 × 2^difficulty, so the chance of hitting it is about e^-32.
  • A difficulty above 26 is refused up front as a misconfiguration. The server's highest tier is 23.

drain_drops_row_after_max_attempts

  • Cause: the test raced the server's own outbox drain, which ticks every 10s. When that tick claimed the due row first, the test's drain_once found nothing, and at the assertion the row was still leased mid-delivery.
  • Fix: new TestApp::drain_outbox(forwarder_id). It makes the forwarder's rows due, drives one pass, then waits until each row has been attempted (delivered, dropped, or rescheduled with one more attempt), whichever drain claimed it.
  • All four outbox call sites use it, including the redelivery test in webhook_signature.rs.
  • The outbox tests now also reach the loopback receiver. Before, their 500s came from the SSRF guard refusing loopback, not from the receiver.

Test plan

  • fmt, clippy (-D warnings)
  • webhook_outbox + webhook_signature ×5, auth_pow ×3: all green
  • A throwaway probe with a rival drain claiming the row the instant it was due: the row was retired and the receiver saw exactly one delivery

🤖 Generated with Claude Code

`successful_login_decays_subnet_failure_counter` failed about one run in
a hundred. The test client's PoW grinder gave up after 10M nonces, and at
difficulty 21 (mean 2^21 tries, geometric) that cap is hit with
probability e^-4.77. The cap is now 32 times the mean, and a difficulty
above 26 is refused up front as a misconfiguration.

`drain_drops_row_after_max_attempts` raced the server's own outbox drain,
which ticks every 10s: when the tick claimed the due row first, the
test's pass found nothing and the row was still leased at the assertion.
`TestApp::drain_outbox` makes a forwarder's rows due, drives a pass and
waits until each row has been attempted, whichever drain claimed it. All
four outbox tests use it. The outbox tests also reach the loopback
receiver now; before, their 500s came from the SSRF guard refusing it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@fylorn
fylorn merged commit 6482690 into dev Sep 24, 2026
@fylorn
fylorn deleted the fix/flaky-integration-tests branch September 24, 2026 05:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant