[Backport release-1.6] fix(e2e): budget tenant scheduling separately from workload readiness - #4435
Conversation
The tenant backend Deployment had a single 300s budget to reach condition=Available, and two unrelated variable costs shared it. The first is scheduling. What the suite establishes before that point is two tenant nodes Ready, which is weaker than schedulable: a Ready node still carries node.cilium.io/agent-not-ready until the tenant cilium agent claims it, and a node the bringup has not finished with is SchedulingDisabled. Scheduling took 2m18s and 1m57s in the two tenant suites of one run, leaving the image pull to finish inside what was left. It did not, and both suites failed with the same message for two different shortfalls. Wait for a node that actually accepts a toleration-free Pod before creating the workload, on its own budget and its own failure message. The gate encodes the scheduler's rule for such a Pod (Ready, not unschedulable, no NoSchedule or NoExecute taint) over a custom-columns probe, prints the node table on both outcomes so a timeout names the taint that held it, and treats a failed probe as not-schedulable so an API blip cannot release it. The readiness wait keeps its 300s, now starting from a schedulable node, and dumps deployment, pod and event state when it runs out. Pin the workload image by digest. The tenant workers reach no registry mirror, so nginx is pulled from Docker Hub on every run either way; the digest fixes what that pull returns instead of leaving a floating tag free to change size and content under a fixed deadline. Assisted-By: Claude <[email protected]> Signed-off-by: Aleksei Sviridkin <[email protected]> (cherry picked from commit 5542ae5) Signed-off-by: Myasnikov Daniil <[email protected]>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: cozystack/cozystack/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ringup writes to the repo root (#4438) Backport of #3891 to `release-1.6`. `hack/e2e-prepare-cluster.bats` on this branch writes the Talos secrets bundle, machine configs, kubeconfig, boot image and `srv*/` dirs to the repo root when the sandbox is brought up outside CI, and nothing in `.gitignore` matches them, so `git add -A` after a local e2e run picks up `secrets.yaml` with the cluster private keys. `hack/run-kubernetes-schedulable_test.bats` from #4435 also leaves `tenantkubeconfig-test-latest-version` in the root on every `make unit-tests`. Cherry-picked clean with `-x`, `.gitignore` only. ```release-note NONE ```
The audit started from labels and dropped every backport PR whose original was not a candidate for the branch. On release-1.6 that hid nine hand backports of unlabelled main PRs (#4431 to #4435, #4437, #4438, #4456 and #4475), and it had no way to notice two open backport PRs for the same originals: #4421 sat open next to #4456 for #3938 and #4280, and nothing reported it. Read the PRs on the branch from the side of the originals they claim as well, and report two more sections per branch. UNLABELLED lists each original claimed by backport PRs on the branch that is not a candidate for it, with every backport PR claiming it and its state, and says whether they backported it, only claim it with a PR still open, or were all closed. It never moves the exit code: the gate answers whether everything labelled landed, and an unlabelled backport can only add to a branch, never leave a labelled change off it. DUPLICATE lists each original claimed by two open backport PRs, or by an open one after another already merged, labelled or not. Both make the audit exit 1. One of the PRs is redundant, or the open one is the rest of a split backport; either way someone has to decide before the cut, and no verdict can say so, since a verdict settles on the first merged backport or reports the first open one as pending. A closed PR next to an open one is how a conflicting bot backport gets redone by hand and is not flagged. The titles of originals that no listing carries come from one GraphQL request for the whole run. A failed lookup costs the titles and nothing else: the URL is derived locally and the exit code is already settled. --json now emits an object per branch holding candidates, unlabelled and duplicates arrays, in place of the bare array of verdicts. On release-1.6 today this lists 13 unlabelled originals, the nine above among them, and three duplicates that are open right now: the bot's conflict drafts for #3936, #4254 and #4292 were left open next to their hand backports, and #4254 and #4292 each also have a fork PR open next to its reopening from a branch in this repository. #4421 is closed and shows up only as a closed claim on #4280. Assisted-by: LLM Signed-off-by: Myasnikov Daniil <[email protected]>
Backport of #3579 to
release-1.6.Tenant backend Deployment in the kubernetes suites still has one 300s budget for scheduling and image pull on this branch. Right now those suites on 1.6 mostly fail earlier, on node-join, so this starts to matter once #3455 widens that window.
Cherry-picked clean with
-x.Testing
make unit-testsgreen, POSIX sh sweep clean.