Repository navigation
Conversation
The proxy marks a job running only when the engine's first response bytes arrive, so a non-streaming job reads queued until it completes. Telling an engine that has started a request from one holding it for a free slot needs the number of requests the engine works on at once. This commit carries that count from engine-manager to the proxy and changes no behavior yet. A manifest's new slots block declares the count. Ollama reads OLLAMA_NUM_PARALLEL from the environment of a process engine-manager launched, and reports 1 when the variable is unset or the engine was adopted. LM Studio reads each loaded instance's config.parallel from the loaded_models response the loaded-model watcher already fetches, so its API is polled no more often. EngineStatus carries the counts as slots while the engine runs, and they are cleared when the run ends. The broker caches the counts from the engine:status read its advertise poll already makes, and sends them on node/set-local-backend while the engine is healthy. They decode apart from the port, so a malformed count costs only the counts, never the engine's registration. The proxy drops counts below 1, keys them by normalized model name and stores them on the facade. Nothing reads them yet. Signed-off-by: Chris Kelsey <[email protected]>
A job read queued until its engine's first response bytes arrived. A non-streaming job on this node's own engine therefore read queued until it finished, and a job waiting for a free slot in the engine read the same as one the engine was processing. With the slot counts the previous commit relays, the proxy can tell the two apart for the requests it sends its own engine. Each facade keeps a slot tracker: for each normalized model, a first-in, first-out queue of the requests it has sent to the local engine. As many as the model's count run, and the rest wait and run in arrival order as earlier ones finish. A request the engine has produced output for runs whatever the count says, and a running request never goes back to waiting. A job is marked running when its ticket runs, and the commit point still marks it running if no slot has. Attempts are now numbered. Each dispatch returns a running job to queued on its next node with no start time, and a slot report from an attempt the request has moved on from changes nothing. The broker store applies a job's events by seq and the desktop shows the latest one, so the return is shown as sent. The tracker decides only what a job's state reads. The scheduler counts queued and running alike as pending, and nothing that routes, retries, reserves or times out a request reads it. A request served by another node still reads running from its commit point. The proxy spec gains section 5.9, and the architecture and known-issues pages say what the states mean and where they are estimates. Signed-off-by: Chris Kelsey <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
A job reads In flight (state
queued) until it reads Running. Today the proxy marks a job running when the engine's first response bytes arrive. A non-streaming job therefore reads In flight until it completes, and a job waiting for a free slot inside the engine reads the same as one the engine is working on. This pull request moves the change to the moment the engine starts the job.slots. Ollama readsOLLAMA_NUM_PARALLELfrom the environment of a process engine-manager launched, and reports 1 when the variable is unset or the engine was adopted. LM Studio reads each loaded instance'sconfig.parallelfrom theloaded_modelsresponse the loaded-model watcher already fetches, so its API is polled no more often.EngineStatuscarries the counts asslotswhile the engine runs. The broker caches them from theengine:statusread its advertise poll already makes and sends them to the proxy onnode/set-local-backend.queuedon that node with no start time, and a late report from the attempt it left changes nothing. The broker store applies a job's events inseqorder (#142), so the return is applied like any other event.Routing does not change. The scheduler counts
queuedandrunningalike as pending work, and nothing that routes, retries, reserves or times out a request reads the slot tracker. The states are estimates, anddocs/known-issues.mdxlists where they can be wrong.This is the third of three pull requests that make job states accurate. It is stacked on #142, and its docs use the In flight label from #141, so it merges after both. Until #142 merges, this pull request targets #142's branch, so CI and the release-intent check do not run on it yet.
Still to come in this draft: a job served by another node still reads Running from its first output, because only that node's proxy can see its engine's slots. A last commit will let that node report when such a request holds a slot, and will update the changelog and docs to match.
Release intent
Changelog title
Jobs show Running as soon as an engine starts them
Changelog body
Bumps
The
nvpair-engine-manager,nvpair-ui-brokerandnvpair-proxybinaries change. The workload manager changes only in its spec.Scope
Audited:
nvpair-engine-manager, which reports the counts. Validation of the new manifestslotsblock, the bundled Ollama and LM Studio manifests and overrides of them, managed launch (the count comes from the launch environment), adoption and command mode (the manifest default), the loaded-model watcher (per-model counts, discarded when they come from an earlier run), andEngineStatus(slotsonly while running, cleared on stop, crash and loss of presence).nvpair-ui-broker, which relays them. The advertise poll caches the counts fromengine:statusand sends them onnode/set-local-backendwhile the engine is healthy. A stopped engine clears the cache and an unavailable engine-manager leaves it alone. The counts decode apart from the port, so a malformed count costs only the counts, never the engine's registration.nvpair-proxy, which uses them.setLocalBackenddrops counts below 1, keys them by normalized model name, and hands them to the tracker after releasingbackendMu. The handler takes a slot ticket just before it sends a request to this node's engine and releases it once the response has been copied, including when the copy aborts. Retry, failover, reservations, the 120-second header and first-content caps, cancellation and the disconnect watcher are unchanged, and every existing proxy test passes unchanged.runningfollowed byqueued. The broker store applies a job's events inseqorder, including after the latest Keep same-numbered jobs apart and apply each job's events in order #142 change to sweep guesses. The workload manager relays them with no state check. The scheduler counts both states as pending on the job's node, so the pending work it ranks by is unchanged. The Electron bridge and the renderer store show the latest event.nvpair-tuihas no workload state logic.EngineStatusgains an optionalslotsfield andnode/set-local-backendan optionalslotsparameter. No consumer decodes either strictly, and a proxy that receives no counts assumes one slot per model. No JSON-RPC method is added or removed, and the generatedservices-api.mdis unchanged.Excluded adjacent work:
services-parity.mdrecords that the counts exist.Validation
Go:
services/nvpair-proxy:go vet ./...andgo test ./...pass, andgo test -race ./...passes in agolang:1.25Linux container. New: 12 tests inslots_test.go(count normalization and the countssetLocalBackendhands the tracker, then the tracker's capacity, per-model queues, promotion order, count changes, output proving a slot, release, model-name normalization, a nil ticket, and concurrent use), 3 injobevents_test.go(a re-point from running, the attempt check, a commit after a slot), and 4 handler tests inlocalslots_test.gorun against both engines (an attempt waiting for a slot, a stream keeping its slot's start time, a failover returning the job to queued, a job cancelled while it waits).services/nvpair-engine-manager:go vet ./...andgo test ./...pass. New: 16 tests inslots_test.go(manifest validation, bundled and overridden manifests, the launch environment, loaded-model parsing and recording, managed, adopted and slot-less engines, and ended runs and unadoption clearing counts) andTestE2EStatusReportsSlots.services/nvpair-ui-broker:go vet ./...andgo test ./...pass. New: 4 advertiser tests for the slot cache and relay.services/nvpair-workload-managerandservices/nvpair-job-scheduler: no code change;go test ./...passes.services/tests(cross-process): CI's command,go test ./... -count=1 -timeout=20m, passes in agolang:1.25Linux container. All 87 tests and subtests ran: 86 passed and 1 skipped (TestModelInventoryRefusesLANPlaintext, because the container has no LAN address). On Windows, this machine's antivirus blocks the unsigned broker binary the suite builds into%TEMP%and stops ago testrun partway, so the Linux run stands in for it.From
desktop/, all passing:npm run typechecknpm run lintnpm run dead-code:checknpm run test:unit(247 passed, 2 skipped: Unix-only wipe-script tests)npm run service-contracts:checknpm run build:modular-binaries -- --forcenode scripts/spdx-headers.mjsreports no missing headers.Still to do by hand:
OLLAMA_NUM_PARALLELunset, so PAIR counts one slot per model, and send two non-streaming requests to one model at once through the local proxy. Confirm one job reads Running and the other In flight until the first finishes.Risk
-racepasses over the proxy suite.workload:submittedafterworkload:started. Every consumer applies events inseqorder or shows the latest one (see Scope).docs/known-issues.mdxlists these.Checklist
git commit -s), certifying the Developer Certificate of Origin.services/versions.jsonis written by automation — do not edit it by hand.