Repository navigation
Conversation
|
Thank you for your PR and interest in PAIR! In general we prefer pulls to pushes, so making the update in |
|
Happy to move it — but I can't find I checked The closest thing to an "apply info" step in the scanner is the tail of node.GPUs = info.GPUs
node.CPU = info.CPU
node.Memory = info.Memoryreached through One thing worth separating out in the push/pull framing, in case it changes the answer: If so, the narrower change is to re-run only the apply step for our own entry: fetch node-info over loopback on the refresh tick and update the existing directory node's For context on why self is the odd one out: If |
|
@Noah-Tervalon-Nvidia — while you are here, this one is still waiting on a question from 12 Sep. I could not find If it lives on an unpublished branch, just say so and I will match its shape once it lands. If you meant the apply step at the tail of Happy to rewrite it either way. I would just rather ask than guess. |
|
@canja006 Sorry for the slow reply and the confusion. |
The local node's hardware figures were captured once and then never refreshed. onBrowse deliberately skips our own entry because self is registry-driven, and publishSelf is otherwise only called when the registry moves: a service (un)register, an identity change, an address change. refreshPeersLoop refreshes advertised addresses, the mesh, models and cluster identity, none of which touch the local figures. Nothing was left to converge them. Measured on an RTX 3060: a model was loaded and the card went to 9.1 GB while the directory kept serving the 0.12 GB it held at daemon start. node-info was serving the true value every two seconds the whole time. refreshSelfInfoOnce pulls the same loopback node-info a peer's browse would, and moves three fields through a new directory.applyInfo that mirrors applyModels: it guards on the node-info endpoint, compares before writing, and reports whether anything actually changed. The record is therefore republished once per real change rather than once per tick, and the whole self record is not rebuilt on a timer. Enrichment dials loopback, never the advertised ip=, keeping the property publishSelfLocked documents: a firewall block on inbound to our own LAN address must not blank the local card. applyInfo treats a nil facet and a zeroed one as different. nil means node-info has never answered for it, which is not the same statement as a reading of zero, and collapsing the two would let a node that stopped reporting look like one reporting idle hardware. Signed-off-by: canja006 <[email protected]>
0b8d8aa to
42f776e
Compare
|
Thanks — that settles it. Rewritten to the narrower shape. What changed since 0b8d8aa:
One deliberate call worth flagging, in case you'd rather it went the other way: Tests in
|
Fixes #4.
onBrowsedeliberately skips our own entry, so the path that re-enriches every peer on every browse never runs for self.publishSelf— the only thing that enriches self — is called at startup, on service (un)register, and after an identity or address change. Nothing was left to converge the local node's GPU, CPU and memory figures, so they stayed at whatever they were when the daemon started.This calls
publishSelf()on therefreshPeersLooptick that already refreshes advertised addresses, the mesh, cluster identity and models.Verified on hardware, two paired nodes (Ubuntu + RTX 3060, macOS):
discovery:get-nodesfollows within 20 s, where before it never moved at allnode-infoand the desktop UI were correct throughout; only the discovery snapshot was frozengo vetandgo test ./...clean inservices/nvpair-node-scanner.One tradeoff worth flagging:
publishSelfemitsnode-updatedunconditionally, so this adds one self update per 15 s tick. The tidier alternative is anapplyInfopath with achangedcheck, mirroringapplyModels— happy to switch to that shape if you prefer it.