[agent] Filed by the scheduled architecture audit routine (CLI and core). Register: discussion #560 register.
Kind: bug. Source: §1 #2; Part 7.2; register C02.
Problem
Both reqwest clients in ApiClient are built with no timeout and no connect_timeout:
None of these paths add a per-request bound:
- the JSON retry loop
send_json_request (#L475-L551), used by fetch_patch, the batch search and org resolution;
fetch_binary for blobs and diffs (#L1061-L1120).
The retry loop retries on status codes only, and a stalled connection never produces one.
Every other HTTP path in the crate is bounded, so the rule exists but is applied per call site:
Reproduced twice on main @ 1169ae6 with an integration test against ApiClient (not committed). The local TCP server accepts the connection, reads the request and never answers. Both fetch_blob(<hash>) and fetch_patch(<uuid>) were still pending when the test's 120 s outer guard fired, on both runs.
Symptoms
None filed. Several ci-janitor flake PRs (#419, #448, #565) handle API blips in tests, but a stall in production has no bound at all.
Impact
scan, get, apply (blob/diff fetch) and vex can hang a CI job until the job-level timeout on a stalled proxy, load balancer or half-open connection. This is P1 and the fix is small.
Proposed change
- Give both clients a
connect_timeout (e.g. 10 s) and a read/idle bound: reqwest's read_timeout, or a per-attempt timeout sized for the request kind. Keep blob/diff bodies on a longer bound than JSON.
- Make the bound one named policy beside
ApiRetryPolicy in api/retry.rs, overridable the same way. A stalled attempt then surfaces as ApiError::Network, which the existing proxy fallback already handles.
- Out of scope here: adding retry to
fetch_binary and merging the three retry systems (C15, a separate refactor).
Size and scope
api/client.rs and api/retry.rs, ~30–60 production lines plus tests.
Acceptance criteria
Dependencies
Blocks nothing. Related to C15 (one retry + timeout primitive) and to #571 (the uncapped blob/diff body in the same fetch_binary).
[agent] Filed by the scheduled architecture audit routine (CLI and core). Register: discussion #560 register.
Kind: bug. Source: §1 #2; Part 7.2; register C02.
Problem
Both reqwest clients in
ApiClientare built with notimeoutand noconnect_timeout:ApiClient::new:api/client.rs#L393-L396plain_client():api/client.rs#L1912-L1922None of these paths add a per-request bound:
send_json_request(#L475-L551), used byfetch_patch, the batch search and org resolution;fetch_binaryfor blobs and diffs (#L1061-L1120).The retry loop retries on status codes only, and a stalled connection never produces one.
Every other HTTP path in the crate is bounded, so the rule exists but is applied per call site:
.timeout(self.vendor_retry.attempt_timeout)at#L1512-L1530, andtokio::time::timeoutat#L1635;telemetry.rs#L296-L298);update/download.rs#L65-L68,update/release.rs#L266-L269);Reproduced twice on main @
1169ae6with an integration test againstApiClient(not committed). The local TCP server accepts the connection, reads the request and never answers. Bothfetch_blob(<hash>)andfetch_patch(<uuid>)were still pending when the test's 120 s outer guard fired, on both runs.Symptoms
None filed. Several
ci-janitorflake PRs (#419, #448, #565) handle API blips in tests, but a stall in production has no bound at all.Impact
scan,get,apply(blob/diff fetch) andvexcan hang a CI job until the job-level timeout on a stalled proxy, load balancer or half-open connection. This is P1 and the fix is small.Proposed change
connect_timeout(e.g. 10 s) and a read/idle bound: reqwest'sread_timeout, or a per-attempttimeoutsized for the request kind. Keep blob/diff bodies on a longer bound than JSON.ApiRetryPolicyinapi/retry.rs, overridable the same way. A stalled attempt then surfaces asApiError::Network, which the existing proxy fallback already handles.fetch_binaryand merging the three retry systems (C15, a separate refactor).Size and scope
api/client.rsandapi/retry.rs, ~30–60 production lines plus tests.Acceptance criteria
fetch_patch, the batch search andfetch_blobeach returnApiError::Networkwithin the configured bound, using a short test override rather than real minutes.CLI_CONTRACT.mdif an env override is added.api_retry_e2e,binary_fetch_error_classification_e2eandblob_fetcher_edges_e2estay green.Dependencies
Blocks nothing. Related to C15 (one retry + timeout primitive) and to #571 (the uncapped blob/diff body in the same
fetch_binary).