Skip to content

test(e2e): raise tenant Kubernetes node-join timeouts to 20m for LINSTOR cold path - #3207

Closed
IvanHunters wants to merge 1 commit into
mainfrom
fix/e2e-node-join-mhc-timeout
Closed

IvanHunters wants to merge 1 commit into
mainfrom
fix/e2e-node-join-mhc-timeout

Conversation

@IvanHunters

Copy link
Copy Markdown
Collaborator

What this PR does

E2E-tenant only: raise the MachineHealthCheck nodeStartupTimeout and the outer >=2 nodes Ready poll deadline in hack/e2e-apps/run-kubernetes.sh from 10m / 12m to 20m / 20m so back-to-back tenant Kubernetes tests do not fail on LINSTOR-CSI cold-provisioning after the preceding tenant teardown. The chart default in packages/apps/kubernetes/values.yaml stays at 10m; production tenants land on Talos images cached on the storage nodes and hit the LINSTOR pool warm, so the tighter budget is correct there.

Root cause

Cozyreport from PR #3044 E2E run 28737984794:

  • Test starts 11:56:29. MachineDeployment fans out both workers at 11:57:06 (xkk94, 5fd5f).
  • 5fd5f: LINSTOR pool had free capacity for its DRBD placement, target PVC bound in ~40s, DV import finished ~12:01, node joined and Ready by ~12:02.
  • xkk94: sat behind kubernetes-latest teardown for pool detach. Target PVC (replicated, autoPlace=3, DRBD 3 replicas) took ~9 minutes to bind. CDI importer only reached Started container importer at 12:07:20, one minute before MHC's 10m nodeStartupTimeout expired at 12:07:06.
  • 12:08:19 — MHC marks xkk94 unhealthy, importer killed mid-run, MachineSet spawns replacement z6c92 at 12:08:21 — 70 seconds before the outer 12m deadline at 12:08:29. Fails on a self-inflicted timeout with no cluster-side fault.

MHC 10m < LINSTOR cold-path (~9m) + CDI import (~3m) + VM boot / kubelet / CNI (~2m) = worker never gets a chance. Raise both budgets in lockstep so the two do not race.

Scope

  • hack/e2e-apps/run-kubernetes.sh only. 28 insertions / 5 deletions.
  • Chart values (packages/apps/kubernetes/values.yaml), CR schema (packages/system/kubernetes-rd/cozyrds/kubernetes.yaml), and the MHC template (packages/apps/kubernetes/templates/cluster.yaml) are unchanged.
  • No product-side timing changes. No new dependencies. No test fixtures touched.

Related

Follows #3206 (bumped sandbox tenant storage quota 100Gi -> 200Gi). Both changes cover the same class of E2E flakes surfaced by back-to-back tenant Kubernetes tests; this one addresses the timing knob after the storage knob.

Release note

NONE

… path

`kubernetes-previous.bats` intermittently fails at "Create a tenant
Kubernetes control plane with previous version" on the sandbox after
`kubernetes-latest.bats` teardown, with only one of the two worker
nodes reaching Ready inside the outer 12m budget. The remaining
Machine gets marked unhealthy by MachineHealthCheck (10m
nodeStartupTimeout) and a replacement is created ~10m after the
Machines' original creation timestamp, leaving no room for it to
finish its own bringup before the outer 12m deadline.

Cozyreport artefact from a run on PR #3044 (28737984794) with all
timing captured:

- Test starts 11:56:29.
- MachineDeployment fans out two Machines at 11:57:06 (xkk94, 5fd5f).
- 5fd5f: replicated 20Gi PVC provisioned in ~40s (LINSTOR CSI already
  had free pool capacity for one placement), DV import finished
  ~12:01, VM Running, node joined and Ready by ~12:02.
- xkk94: sat behind kubernetes-latest teardown for replicated pool
  detach; target PVC (autoPlace=3, DRBD) took ~9 minutes to bind, so
  the CDI importer only reached `Started container importer` at
  12:07:20, one minute before MHC's 10m nodeStartupTimeout expired at
  12:07:06. At 12:08:19 MHC marks xkk94 unhealthy, the importer is
  killed mid-run, MachineSet creates the replacement z6c92 at
  12:08:21 — 70 seconds before the outer 12m deadline at 12:08:29.
  Fails on a self-inflicted timeout with no cluster-side fault.

Raise both budgets in the e2e Kubernetes CR spec to 20m:

- `nodeHealthCheck.nodeStartupTimeout: "20m"` on the tenant CR so MHC
  waits for the LINSTOR + CDI + Talos + kubelet + CNI chain instead
  of remediating a still-provisioning worker.
- Outer bats deadline `timeout 12m` -> `timeout 20m` on the
  `>=2 nodes Ready` poll so the test grants the same window it just
  asked MHC to observe.

E2E-tenant scope only. The chart default in
packages/apps/kubernetes/values.yaml stays at 10m -- production
tenants land on Talos images cached on the storage nodes and hit the
LINSTOR pool warm, so the tighter budget is correct for them and
only the sandbox's back-to-back tenant Kubernetes tests need the
slack.

Signed-off-by: Ivan Okhotnikov <[email protected]>
@coderabbitai

coderabbitai Bot commented Jul 5, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: cfb5babd-3036-450f-bef9-86aa27b1118c

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/e2e-node-join-mhc-timeout

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size/M This PR changes 30-99 lines, ignoring generated files area/testing Issues or PRs related to testing (e2e, bats, unit tests) labels Jul 5, 2026
@IvanHunters IvanHunters closed this Jul 5, 2026
@IvanHunters

Copy link
Copy Markdown
Collaborator Author

Closing — this was a workaround for a symptom, not the root cause.

Root cause is proven and lives in LINSTOR-CSI v1.10.6 CreateVolume path, not in the E2E test infrastructure. From linstor-csi-controller log in the cozyreport artefact of run 28737984794 (job 85216039506), for PV pvc-a3cd828e-1430-471a-aab2-f1862db1fbdc (xkk94 worker):

CreateVolume failed for pvc-a3cd828e-...: rpc error: code = ResourceExhausted
  desc = failed to enough replicas on requisite nodes:
  Message: 'Not enough available nodes';
  Details: 'Not enough nodes fulfilling the following auto-place criteria:
    * has a deployed storage pool named [data]
    * the storage pools have to have at least '20971520' free space
    * ...
    Replica count: 3
    Additional replica count: 3
    Node names: srv1, srv2, srv3
    Do not place with resource: pvc-a3cd828e-...   <-- self-anti-affinity
    Layer stack: DRBD, STORAGE'

LINSTOR-CSI retries CreateVolume on the same PV name; each retry adds an anti-affinity constraint against the previous still-being-deleted placement of the SAME PV name (Do not place with resource: pvc-<same-uuid>), which pins the placement to zero remaining nodes and returns ResourceExhausted. The driver then deletes and re-tries. Ten create/delete cycles later (11:57:09 to 12:06:14 in the trace) the timing race resolves and the PVC binds.

The linstor-satellite side confirms this from the srv3 satellite log:

2026-07-05 11:57:24.046 ERROR LINSTOR/Satellite/8a3a83 SYSTEM - Failed to create zfsvolume [Report number 6A4A394E-E4EA0-000000]
2026-07-05 11:57:24.308 ERROR LINSTOR/Satellite/924026 SYSTEM - Failed to create zfsvolume [Report number 6A4A394E-E4EA0-000001]

xkk94 accumulates a RetryTask: Failed resource ... of node 'srv3' added for retry — the sibling worker (5fd5f, created in the same second) goes through the same create/delete cycle count (9-10 rounds) but never lands in the anti-affinity dead zone, and its PVC binds in ~5m instead of ~9m. Deterministically random depending on which one enters the retry-task queue first.

The MHC 10m nodeStartupTimeout + outer 12m budget I proposed in this PR do not remove the race, they only raise the ceiling that the race has to hit. The fix belongs upstream in piraeusdatastore/linstor-csi — CreateVolume must not queue an anti-affinity constraint against a placement that is being torn down by the same driver in the same reconciliation cycle. Filing that separately is out of scope for the e2e stabilisation batch.

Not merging this.

@IvanHunters
IvanHunters deleted the fix/e2e-node-join-mhc-timeout branch July 5, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/testing Issues or PRs related to testing (e2e, bats, unit tests) size/M This PR changes 30-99 lines, ignoring generated files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant