test(e2e): raise tenant Kubernetes node-join timeouts to 20m for LINSTOR cold path - #3207
IvanHunters wants to merge 1 commit into
Conversation
… path `kubernetes-previous.bats` intermittently fails at "Create a tenant Kubernetes control plane with previous version" on the sandbox after `kubernetes-latest.bats` teardown, with only one of the two worker nodes reaching Ready inside the outer 12m budget. The remaining Machine gets marked unhealthy by MachineHealthCheck (10m nodeStartupTimeout) and a replacement is created ~10m after the Machines' original creation timestamp, leaving no room for it to finish its own bringup before the outer 12m deadline. Cozyreport artefact from a run on PR #3044 (28737984794) with all timing captured: - Test starts 11:56:29. - MachineDeployment fans out two Machines at 11:57:06 (xkk94, 5fd5f). - 5fd5f: replicated 20Gi PVC provisioned in ~40s (LINSTOR CSI already had free pool capacity for one placement), DV import finished ~12:01, VM Running, node joined and Ready by ~12:02. - xkk94: sat behind kubernetes-latest teardown for replicated pool detach; target PVC (autoPlace=3, DRBD) took ~9 minutes to bind, so the CDI importer only reached `Started container importer` at 12:07:20, one minute before MHC's 10m nodeStartupTimeout expired at 12:07:06. At 12:08:19 MHC marks xkk94 unhealthy, the importer is killed mid-run, MachineSet creates the replacement z6c92 at 12:08:21 — 70 seconds before the outer 12m deadline at 12:08:29. Fails on a self-inflicted timeout with no cluster-side fault. Raise both budgets in the e2e Kubernetes CR spec to 20m: - `nodeHealthCheck.nodeStartupTimeout: "20m"` on the tenant CR so MHC waits for the LINSTOR + CDI + Talos + kubelet + CNI chain instead of remediating a still-provisioning worker. - Outer bats deadline `timeout 12m` -> `timeout 20m` on the `>=2 nodes Ready` poll so the test grants the same window it just asked MHC to observe. E2E-tenant scope only. The chart default in packages/apps/kubernetes/values.yaml stays at 10m -- production tenants land on Talos images cached on the storage nodes and hit the LINSTOR pool warm, so the tighter budget is correct for them and only the sandbox's back-to-back tenant Kubernetes tests need the slack. Signed-off-by: Ivan Okhotnikov <[email protected]>
|
Important Review skippedDraft detected. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Closing — this was a workaround for a symptom, not the root cause. Root cause is proven and lives in LINSTOR-CSI v1.10.6 CreateVolume path, not in the E2E test infrastructure. From LINSTOR-CSI retries CreateVolume on the same PV name; each retry adds an anti-affinity constraint against the previous still-being-deleted placement of the SAME PV name ( The linstor-satellite side confirms this from the srv3 satellite log: xkk94 accumulates a The MHC 10m nodeStartupTimeout + outer 12m budget I proposed in this PR do not remove the race, they only raise the ceiling that the race has to hit. The fix belongs upstream in Not merging this. |
What this PR does
E2E-tenant only: raise the MachineHealthCheck
nodeStartupTimeoutand the outer>=2 nodes Readypoll deadline inhack/e2e-apps/run-kubernetes.shfrom 10m / 12m to 20m / 20m so back-to-back tenant Kubernetes tests do not fail on LINSTOR-CSI cold-provisioning after the preceding tenant teardown. The chart default inpackages/apps/kubernetes/values.yamlstays at 10m; production tenants land on Talos images cached on the storage nodes and hit the LINSTOR pool warm, so the tighter budget is correct there.Root cause
Cozyreport from PR #3044 E2E run 28737984794:
xkk94,5fd5f).5fd5f: LINSTOR pool had free capacity for its DRBD placement, target PVC bound in ~40s, DV import finished ~12:01, node joined and Ready by ~12:02.xkk94: sat behindkubernetes-latestteardown for pool detach. Target PVC (replicated,autoPlace=3, DRBD 3 replicas) took ~9 minutes to bind. CDI importer only reachedStarted container importerat 12:07:20, one minute before MHC's 10mnodeStartupTimeoutexpired at 12:07:06.xkk94unhealthy, importer killed mid-run, MachineSet spawns replacementz6c92at 12:08:21 — 70 seconds before the outer 12m deadline at 12:08:29. Fails on a self-inflicted timeout with no cluster-side fault.MHC 10m < LINSTOR cold-path (~9m) + CDI import (~3m) + VM boot / kubelet / CNI (~2m) = worker never gets a chance. Raise both budgets in lockstep so the two do not race.
Scope
hack/e2e-apps/run-kubernetes.shonly. 28 insertions / 5 deletions.packages/apps/kubernetes/values.yaml), CR schema (packages/system/kubernetes-rd/cozyrds/kubernetes.yaml), and the MHC template (packages/apps/kubernetes/templates/cluster.yaml) are unchanged.Related
Follows #3206 (bumped sandbox tenant storage quota 100Gi -> 200Gi). Both changes cover the same class of E2E flakes surfaced by back-to-back tenant Kubernetes tests; this one addresses the timing knob after the storage knob.
Release note