Skip to content

fix(credentials): do not cool down keys for at-capacity errors - #6

Merged
yelixir-dev merged 1 commit into
yelixir-dev:mainfrom
A8Cl233395:fix/at-capacity
Oct 2, 2026
Merged

yelixir-dev merged 1 commit into
yelixir-dev:mainfrom
A8Cl233395:fix/at-capacity

Conversation

@A8Cl233395

Copy link
Copy Markdown
Contributor

Fix for the at-capacity cooldown incident on the single-key deployment. Any terminal upstream failure used to call recordFailure, including at-capacity errors that arrive as statusless SSE error events and transient failures that are retried and eventually succeed. With one credential, that benched the only key for 60s and every request failed in ~1ms with NoAvailableCommandCodeCredentialError until the window expired.

Changes

src/commandcode.ts (alpha /alpha/generate path)

  • Only credential-scoped statuses (401/402/403, isFatalCredFailure) record a failure and start a cooldown; HTTP non-ok 429/5xx responses, empty bodies, stream error events and network errors now call finalizeRelease() — releasing the in-flight slot without touching disabledUntil.
  • Statusless upstream stream error events are never recorded as credential faults. They are still retried while no visible event has been emitted, then forwarded to the caller as-is.
  • Renamed finalizeCallerAbort to finalizeRelease and routed the abort path through it, so caller aborts no longer leave a cooldown.
  • Attempt zero: when the router reports NoAvailableCommandCodeCredentialError, selection is retried with ignoreCooldown: true so a cooling credential can still serve instead of failing the request.

src/provider.ts (provider /provider/v1/chat/completions path)

  • Same release policy: added finalizeRelease and made both the !response.ok branch and the thrown-error branch record a failure only for fatal credential statuses (401/402/403); everything else releases.
  • Fatal credentials are still excluded from further attempts within the request.

tests/commandcode.test.ts

  • Added four regression tests: retry-then-success leaves no stale cooldown; a statusless at-capacity error does not cool the key; exhausted 429 retries do not cool the key; an injected cooldown is served through.
  • No wire-format or API-shape changes; existing failover, opaque-retryable-error and abort tests are unchanged.

Resulting behavior

An at-capacity response now retries within the request (default budget 5 attempts, ~3.75s of backoff worst case). If it still fails, or if visible output was already streamed, the upstream error is forwarded to the caller — an in-band SSE error frame (commandcode_event_error) for streaming, HTTP 502 JSON for non-streaming — and no cooldown is recorded, so the next request is served normally.

Tests

  • tsc --noEmit, eslint and the build pass.
  • Full vitest suite: 265 passed, 4 pre-existing Windows-only failures (path separators, 0600 permissions, executable bits, launchctl ENOENT).

Any terminal upstream failure used to record a 60s credential cooldown,
including capacity errors that arrive as statusless SSE error events and
transient blips that are retried and eventually succeed. With a single
key that benched the only credential, so every request for the next
minute failed instantly with NoAvailableCommandCodeCredentialError.

Only credential-scoped failures (401/402/403) may start a cooldown now.
Provider-scoped failures -- statusless stream error events, 429/5xx
responses, empty bodies, and network errors -- release the credential
without marking it, while still being retried within the request up to
the configured attempt budget. Cooldowns are treated as a preference
rather than a gate: attempt zero falls back to an ignoreCooldown
selection, so a cooling credential can still serve instead of failing
the request. The provider client follows the same release policy.

Cover the CommandCode client with regression cases for retry-then-success
without a stale cooldown, statusless at-capacity errors, exhausted 429
retries, and serving through an injected cooldown.
@yelixir-dev

Copy link
Copy Markdown
Owner

Thanks for tracking this down and for the clear write-up — benching the only key for 60s on a transient at-capacity error was a real outage mode for single-key deployments.

Verified on our side before merging: rebased onto current main (which gained CLI 1.66.0 alignment, the new dashboard and session affinity since you opened this) with no conflicts; tsc --noEmit, eslint and the full vitest suite (16 files, 284 tests) pass on Linux, and CI is green on Node 20/22/24.

Merging as-is. We'll follow up on main with a release commit that updates the docs that still describe 429/5xx/timeout cooldowns, and bump to 1.66.0.d.

@yelixir-dev
yelixir-dev merged commit ec97dcd into yelixir-dev:main Oct 2, 2026
3 checks passed
yelixir-dev added a commit that referenced this pull request Oct 2, 2026
Release PR #6 (do not cool down keys for provider-scoped failures):
429/5xx, empty bodies, stream errors and network errors now retry
within the request without benching the credential, so a single-key
deployment no longer fails every request for 60s after one
at-capacity error.

Update the README routing paragraph (en/ko/zh), KNOW_HOW,
ARCHITECTURE and both deployment guides, which still described
429/5xx/timeout cooldowns, and record the release in PROCESS_LOG.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants