Skip to content

[Backport release-1.6] fix(e2e): budget tenant scheduling separately from workload readiness - #4435

Merged
myasnikovdaniil merged 1 commit into
release-1.6from
backport-3579-to-release-1.6
Sep 24, 2026
Merged

myasnikovdaniil merged 1 commit into
release-1.6from
backport-3579-to-release-1.6

Conversation

@myasnikovdaniil

Copy link
Copy Markdown
Contributor

Backport of #3579 to release-1.6.

Tenant backend Deployment in the kubernetes suites still has one 300s budget for scheduling and image pull on this branch. Right now those suites on 1.6 mostly fail earlier, on node-join, so this starts to matter once #3455 widens that window.

Cherry-picked clean with -x.

Testing

  • make unit-tests green, POSIX sh sweep clean.
NONE

The tenant backend Deployment had a single 300s budget to reach
condition=Available, and two unrelated variable costs shared it. The
first is scheduling. What the suite establishes before that point is two
tenant nodes Ready, which is weaker than schedulable: a Ready node still
carries node.cilium.io/agent-not-ready until the tenant cilium agent
claims it, and a node the bringup has not finished with is
SchedulingDisabled. Scheduling took 2m18s and 1m57s in the two tenant
suites of one run, leaving the image pull to finish inside what was
left. It did not, and both suites failed with the same message for two
different shortfalls.

Wait for a node that actually accepts a toleration-free Pod before
creating the workload, on its own budget and its own failure message.
The gate encodes the scheduler's rule for such a Pod (Ready, not
unschedulable, no NoSchedule or NoExecute taint) over a custom-columns
probe, prints the node table on both outcomes so a timeout names the
taint that held it, and treats a failed probe as not-schedulable so an
API blip cannot release it. The readiness wait keeps its 300s, now
starting from a schedulable node, and dumps deployment, pod and event
state when it runs out.

Pin the workload image by digest. The tenant workers reach no registry
mirror, so nginx is pulled from Docker Hub on every run either way; the
digest fixes what that pull returns instead of leaving a floating tag
free to change size and content under a fixed deadline.

Assisted-By: Claude <[email protected]>
Signed-off-by: Aleksei Sviridkin <[email protected]>
(cherry picked from commit 5542ae5)
Signed-off-by: Myasnikov Daniil <[email protected]>
@coderabbitai

coderabbitai Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: cozystack/cozystack/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 573132e1-1b51-4e73-bfc4-250d394108e8

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added area/release Issues or PRs related to release tooling (changelog, backport, release pipeline) area/testing Issues or PRs related to testing (e2e, bats, unit tests) kind/bug Categorizes issue or PR as related to a bug size/L This PR changes 100-499 lines, ignoring generated files labels Sep 24, 2026
@myasnikovdaniil
myasnikovdaniil merged commit 78f1d2a into release-1.6 Sep 24, 2026
13 of 14 checks passed
@myasnikovdaniil
myasnikovdaniil deleted the backport-3579-to-release-1.6 branch September 24, 2026 09:53
myasnikovdaniil added a commit that referenced this pull request Sep 24, 2026
…ringup writes to the repo root (#4438)

Backport of #3891 to `release-1.6`.

`hack/e2e-prepare-cluster.bats` on this branch writes the Talos secrets
bundle, machine configs, kubeconfig, boot image and `srv*/` dirs to the
repo root when the sandbox is brought up outside CI, and nothing in
`.gitignore` matches them, so `git add -A` after a local e2e run picks
up `secrets.yaml` with the cluster private keys.
`hack/run-kubernetes-schedulable_test.bats` from #4435 also leaves
`tenantkubeconfig-test-latest-version` in the root on every `make
unit-tests`.

Cherry-picked clean with `-x`, `.gitignore` only.

```release-note
NONE
```
myasnikovdaniil added a commit that referenced this pull request Sep 25, 2026
The audit started from labels and dropped every backport PR whose
original was not a candidate for the branch. On release-1.6 that hid
nine hand backports of unlabelled main PRs (#4431 to #4435, #4437,
#4438, #4456 and #4475), and it had no way to notice two open backport
PRs for the same originals: #4421 sat open next to #4456 for #3938 and
#4280, and nothing reported it.

Read the PRs on the branch from the side of the originals they claim as
well, and report two more sections per branch.

UNLABELLED lists each original claimed by backport PRs on the branch
that is not a candidate for it, with every backport PR claiming it and
its state, and says whether they backported it, only claim it with a PR
still open, or were all closed. It never moves the exit code: the gate
answers whether everything labelled landed, and an unlabelled backport
can only add to a branch, never leave a labelled change off it.

DUPLICATE lists each original claimed by two open backport PRs, or by
an open one after another already merged, labelled or not. Both make
the audit exit 1. One of the PRs is redundant, or the open one is the
rest of a split backport; either way someone has to decide before the
cut, and no verdict can say so, since a verdict settles on the first
merged backport or reports the first open one as pending. A closed PR
next to an open one is how a conflicting bot backport gets redone by
hand and is not flagged.

The titles of originals that no listing carries come from one GraphQL
request for the whole run. A failed lookup costs the titles and nothing
else: the URL is derived locally and the exit code is already settled.

--json now emits an object per branch holding candidates, unlabelled
and duplicates arrays, in place of the bare array of verdicts.

On release-1.6 today this lists 13 unlabelled originals, the nine above
among them, and three duplicates that are open right now: the bot's
conflict drafts for #3936, #4254 and #4292 were left open next to their
hand backports, and #4254 and #4292 each also have a fork PR open next
to its reopening from a branch in this repository. #4421 is closed and
shows up only as a closed claim on #4280.

Assisted-by: LLM
Signed-off-by: Myasnikov Daniil <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/release Issues or PRs related to release tooling (changelog, backport, release pipeline) area/testing Issues or PRs related to testing (e2e, bats, unit tests) kind/bug Categorizes issue or PR as related to a bug size/L This PR changes 100-499 lines, ignoring generated files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants