fix(server): send float16 multivectors to the sidecar as bytes - #416
Conversation
On the cluster path the worker hands each encode result to the Rust sidecar over IPC. Multivectors went as lists of Python floats, which msgpack packs as 9-byte doubles, in one response capped at 128 MiB. Wide models overflow it: a full 8,192-token batch of TopK-Embed-small (2,048-dim tokens) packed to 130 MiB and one 8,192-token document to 144 MiB, so the whole batch failed. The sidecar already takes a float16 matrix as little-endian bytes (MultivectorOutput.values_f16) and says so on every encode batch (accepts_batched_f16_multivectors). The worker now sends float16 multivectors that way when the flag is set: 2 bytes a value, so the same batch is 29 MiB and the document 32 MiB. float32 outputs, and sidecars without the flag, keep the list. Tests: bytes only with the flag and only for float16, the outcome through process_encode_batch, and the packed size per value.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. 9 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 10 reviews per hour. 📝 WalkthroughWalkthroughThe encode queue now supports little-endian byte output for float16 multivectors when the request enables the capability. Tests cover byte framing, list output, float32 output, packed size, and end-to-end encoding. ChangesFloat16 multivector encoding
Sequence Diagram(s)sequenceDiagram
participant process_encode_batch
participant _run_encode_group
participant _maybe_multivector_raw_output
process_encode_batch->>_run_encode_group: Pass accepts_batched_f16_multivectors
_run_encode_group->>_maybe_multivector_raw_output: Pass f16_bytes for multivector output
_maybe_multivector_raw_output-->>_run_encode_group: Return values_f16 bytes or list values
Suggested reviewers: Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to The float16 byte output matches the sidecar’s supported format and preserves existing output for other requests. No actionable merge-blocking risk remains; normal test checks should complete before merging. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Summary
In SIE's cluster setup, which managed SIE uses, a request does not reach the Python server over HTTP. It goes from the gateway through the queue to a Rust sidecar that runs next to each GPU worker, and the sidecar talks to the Python worker over a local socket (IPC). When the worker hands encode results back, multi-vector outputs were converted to lists of Python floats, which msgpack packs as 9-byte doubles, and the whole batch's response is capped at 128 MiB.
Models with wide multi-vector outputs overflow that cap, and then the whole batch fails. TopK-Embed-V1, being added in #417, returns 2,048 numbers for every token: one ordinary full batch of it is about 130 MiB on this hop. Wide models SIE already ships can hit the same cap with large enough batches, for example
nvidia/llama-nemoretriever-colembed-3b-v1(3,072 numbers per token) andnvidia/nemotron-colembed-vl-4b-v2(2,560).The sidecar already accepts a float16 matrix as raw little-endian bytes (
MultivectorOutput.values_f16, whichpublisher.rsturns into the wire payload) and says so on every encode batch (accepts_batched_f16_multivectors). The worker now uses that for float16 multivectors: 2 bytes a value instead of 9, so the same batch is about 29 MiB.Sizes
Measured with SIE's own packing code (
_maybe_multivector_raw_output, the response envelope,msgpack.packb), before and after:Compatibility
Only float16 multivectors change, and only when the sidecar sets the flag; production sidecars always do (
dispatcher.rs). float32 outputs, and sidecars without the flag, keep the list of floats. The single-server HTTP path does not use this hop and is unchanged.Testing
test_queue_executor_stage1d.py: bytes are sent only with the flag and only for float16, the outcome throughprocess_encode_batchcarries them, and a packed matrix takes about 2 bytes a value instead of 9.Summary by CodeRabbit