Skip to content

Show a job as running once its engine starts it - #143

Draft
ckelseynv wants to merge 2 commits into
feature/job-identity-and-event-orderfrom
feature/accurate-job-states
Draft

ckelseynv wants to merge 2 commits into
feature/job-identity-and-event-orderfrom
feature/accurate-job-states

Conversation

@ckelseynv

Copy link
Copy Markdown
Collaborator

Description

A job reads In flight (state queued) until it reads Running. Today the proxy marks a job running when the engine's first response bytes arrive. A non-streaming job therefore reads In flight until it completes, and a job waiting for a free slot inside the engine reads the same as one the engine is working on. This pull request moves the change to the moment the engine starts the job.

  • Engines report how many requests they process at once. An engine manifest can now declare slots. Ollama reads OLLAMA_NUM_PARALLEL from the environment of a process engine-manager launched, and reports 1 when the variable is unset or the engine was adopted. LM Studio reads each loaded instance's config.parallel from the loaded_models response the loaded-model watcher already fetches, so its API is polled no more often. EngineStatus carries the counts as slots while the engine runs. The broker caches them from the engine:status read its advertise poll already makes and sends them to the proxy on node/set-local-backend.
  • The proxy tracks the slots of its own node's engine. Each facade keeps, per model, a first-in, first-out queue of the requests it has sent to that engine. As many as the model's count run, and the rest wait their turn. A job reads Running as soon as its request holds a slot, or as soon as the engine produces output for it if that comes first. A running request never goes back to waiting.
  • A failover returns a job to In flight. Attempts are numbered. When a request moves to another node, its job returns to queued on that node with no start time, and a late report from the attempt it left changes nothing. The broker store applies a job's events in seq order (#142), so the return is applied like any other event.

Routing does not change. The scheduler counts queued and running alike as pending work, and nothing that routes, retries, reserves or times out a request reads the slot tracker. The states are estimates, and docs/known-issues.mdx lists where they can be wrong.

This is the third of three pull requests that make job states accurate. It is stacked on #142, and its docs use the In flight label from #141, so it merges after both. Until #142 merges, this pull request targets #142's branch, so CI and the release-intent check do not run on it yet.

Still to come in this draft: a job served by another node still reads Running from its first output, because only that node's proxy can see its engine's slots. A last commit will let that node report when such a request holds a slot, and will update the changelog and docs to match.

Release intent

Changelog title

Jobs show Running as soon as an engine starts them

Changelog body

  • A job served by the engine on the computer that received it shows Running as soon as the engine starts working on it, rather than when the engine first returns output. A job waiting for its turn in the engine shows In flight.
  • PAIR works this out from how many requests each engine runs at once, so it is an estimate. Known Issues lists where it can be wrong.
  • A job that moves to another node after a failure shows In flight again until that node's engine starts it.

Bumps

  • services: minor
  • nvpair-cluster-manager: none
  • nvpair-engine-manager: minor
  • nvpair-errors: none
  • nvpair-job-scheduler: none
  • nvpair-manual-nodes: none
  • nvpair-node-info: none
  • nvpair-node-scanner: none
  • nvpair-node-settings: none
  • nvpair-proxy: minor
  • nvpair-tui: none
  • nvpair-ui-broker: minor
  • nvpair-workload-manager: none

The nvpair-engine-manager, nvpair-ui-broker and nvpair-proxy binaries change. The workload manager changes only in its spec.

Scope

Audited:

  • nvpair-engine-manager, which reports the counts. Validation of the new manifest slots block, the bundled Ollama and LM Studio manifests and overrides of them, managed launch (the count comes from the launch environment), adoption and command mode (the manifest default), the loaded-model watcher (per-model counts, discarded when they come from an earlier run), and EngineStatus (slots only while running, cleared on stop, crash and loss of presence).
  • nvpair-ui-broker, which relays them. The advertise poll caches the counts from engine:status and sends them on node/set-local-backend while the engine is healthy. A stopped engine clears the cache and an unavailable engine-manager leaves it alone. The counts decode apart from the port, so a malformed count costs only the counts, never the engine's registration.
  • nvpair-proxy, which uses them. setLocalBackend drops counts below 1, keys them by normalized model name, and hands them to the tracker after releasing backendMu. The handler takes a slot ticket just before it sends a request to this node's engine and releases it once the response has been copied, including when the copy aborts. Retry, failover, reservations, the 120-second header and first-content caps, cancellation and the disconnect watcher are unchanged, and every existing proxy test passes unchanged.
  • Every consumer that can now see running followed by queued. The broker store applies a job's events in seq order, including after the latest Keep same-numbered jobs apart and apply each job's events in order #142 change to sweep guesses. The workload manager relays them with no state check. The scheduler counts both states as pending on the job's node, so the pending work it ranks by is unchanged. The Electron bridge and the renderer store show the latest event. nvpair-tui has no workload state logic.
  • The wire. EngineStatus gains an optional slots field and node/set-local-backend an optional slots parameter. No consumer decodes either strictly, and a proxy that receives no counts assumes one slot per model. No JSON-RPC method is added or removed, and the generated services-api.md is unchanged.

Excluded adjacent work:

  • Jobs served by another node, which the last commit of this draft covers.
  • Manual nodes tell the proxy nothing about their slots, so their jobs read Running from their first output.
  • Requests sent straight to an engine's own port bypass the proxy and are not counted, so a proxied job can read Running while it waits behind them.
  • The desktop shows no slot counts and needs no code change. services-parity.md records that the counts exist.

Validation

Go:

  • services/nvpair-proxy: go vet ./... and go test ./... pass, and go test -race ./... passes in a golang:1.25 Linux container. New: 12 tests in slots_test.go (count normalization and the counts setLocalBackend hands the tracker, then the tracker's capacity, per-model queues, promotion order, count changes, output proving a slot, release, model-name normalization, a nil ticket, and concurrent use), 3 in jobevents_test.go (a re-point from running, the attempt check, a commit after a slot), and 4 handler tests in localslots_test.go run against both engines (an attempt waiting for a slot, a stream keeping its slot's start time, a failover returning the job to queued, a job cancelled while it waits).
  • services/nvpair-engine-manager: go vet ./... and go test ./... pass. New: 16 tests in slots_test.go (manifest validation, bundled and overridden manifests, the launch environment, loaded-model parsing and recording, managed, adopted and slot-less engines, and ended runs and unadoption clearing counts) and TestE2EStatusReportsSlots.
  • services/nvpair-ui-broker: go vet ./... and go test ./... pass. New: 4 advertiser tests for the slot cache and relay.
  • services/nvpair-workload-manager and services/nvpair-job-scheduler: no code change; go test ./... passes.
  • services/tests (cross-process): CI's command, go test ./... -count=1 -timeout=20m, passes in a golang:1.25 Linux container. All 87 tests and subtests ran: 86 passed and 1 skipped (TestModelInventoryRefusesLANPlaintext, because the container has no LAN address). On Windows, this machine's antivirus blocks the unsigned broker binary the suite builds into %TEMP% and stops a go test run partway, so the Linux run stands in for it.

From desktop/, all passing:

  • npm run typecheck
  • npm run lint
  • npm run dead-code:check
  • npm run test:unit (247 passed, 2 skipped: Unix-only wipe-script tests)
  • npm run service-contracts:check
  • npm run build:modular-binaries -- --force

node scripts/spdx-headers.mjs reports no missing headers.

Still to do by hand:

  • Start Ollama from PAIR with OLLAMA_NUM_PARALLEL unset, so PAIR counts one slot per model, and send two non-streaming requests to one model at once through the local proxy. Confirm one job reads Running and the other In flight until the first finishes.
  • In LM Studio, load a model with parallel set to 2, wait about ten seconds, send three requests at once, and confirm two read Running.
  • On a two-node cluster, stop the local engine while a job runs. Confirm the job fails over, reads In flight on its next node, and then Running there.

Risk

  • Display only. The tracker decides only what a job's state reads. A wrong count makes a state wrong, never a route.
  • Concurrency. The tracker has its own lock, which is never held across a call out; its only side effect is closing a ticket's channel. A waiting local attempt starts one goroutine, which ends when its ticket runs or the attempt ends. -race passes over the proxy suite.
  • Running is no longer final. A failover sends workload:submitted after workload:started. Every consumer applies events in seq order or shows the latest one (see Scope).
  • Estimates. Loading a model counts as running. Ollama's parallel setting is not visible when PAIR did not start Ollama with it. LM Studio's per-model counts arrive up to about ten seconds after a load, and two requests that arrive at the same moment may swap. docs/known-issues.mdx lists these.

Checklist

  • I have read the Contributing Guidelines.
  • Every commit is signed off (git commit -s), certifying the Developer Certificate of Origin.
  • New or existing tests cover the change.
  • Relevant documentation is updated.
  • I checked the diff, changed filenames, and commit messages for credentials, private data, internal URLs, internal issue identifiers, and generated artifacts.
  • I recorded the validation commands and results above.
  • I declared version bumps in the release-intent block above. services/versions.json is written by automation — do not edit it by hand.

The proxy marks a job running only when the engine's first response bytes arrive, so a non-streaming job reads queued until it completes. Telling an engine that has started a request from one holding it for a free slot needs the number of requests the engine works on at once. This commit carries that count from engine-manager to the proxy and changes no behavior yet.

A manifest's new slots block declares the count. Ollama reads OLLAMA_NUM_PARALLEL from the environment of a process engine-manager launched, and reports 1 when the variable is unset or the engine was adopted. LM Studio reads each loaded instance's config.parallel from the loaded_models response the loaded-model watcher already fetches, so its API is polled no more often. EngineStatus carries the counts as slots while the engine runs, and they are cleared when the run ends.

The broker caches the counts from the engine:status read its advertise poll already makes, and sends them on node/set-local-backend while the engine is healthy. They decode apart from the port, so a malformed count costs only the counts, never the engine's registration. The proxy drops counts below 1, keys them by normalized model name and stores them on the facade. Nothing reads them yet.

Signed-off-by: Chris Kelsey <[email protected]>
A job read queued until its engine's first response bytes arrived. A non-streaming job on this node's own engine therefore read queued until it finished, and a job waiting for a free slot in the engine read the same as one the engine was processing. With the slot counts the previous commit relays, the proxy can tell the two apart for the requests it sends its own engine.

Each facade keeps a slot tracker: for each normalized model, a first-in, first-out queue of the requests it has sent to the local engine. As many as the model's count run, and the rest wait and run in arrival order as earlier ones finish. A request the engine has produced output for runs whatever the count says, and a running request never goes back to waiting. A job is marked running when its ticket runs, and the commit point still marks it running if no slot has.

Attempts are now numbered. Each dispatch returns a running job to queued on its next node with no start time, and a slot report from an attempt the request has moved on from changes nothing. The broker store applies a job's events by seq and the desktop shows the latest one, so the return is shown as sent.

The tracker decides only what a job's state reads. The scheduler counts queued and running alike as pending, and nothing that routes, retries, reserves or times out a request reads it. A request served by another node still reads running from its commit point. The proxy spec gains section 5.9, and the architecture and known-issues pages say what the states mean and where they are estimates.

Signed-off-by: Chris Kelsey <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant